MulCogBench: a multi-modal cognitive benchmark dataset for evaluating Chinese and English computational language models.

Pre-trained computational language models have recently made remarkable progress in harnessing the language abilities which were considered unique to humans. Their success has raised interest in whether these models represent and process language like humans. To answer this question, this paper prop...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 3005 - 3029
Autores principales: Zhang, Yunhao, Zhang, Xiaohan, Li, Chong, Wang, Shaonan, Zong, Chengqing
Formato: Conference Paper/Materials
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909096&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909096
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909096
        10.1007/s10579-025-09843-2
      ppf: 3005
      ppct: 24
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 4.3MB
      tig:
        atl: MulCogBench: a multi-modal cognitive benchmark dataset for evaluating Chinese and English computational language models.
      aug:
        au:
          Zhang, Yunhao
          Zhang, Xiaohan
          Li, Chong
          Wang, Shaonan
          Zong, Chengqing
        affil:
          https://ror.org/022c3hy66 State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS, Beijing, China
          https://ror.org/05qbk4x57 School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
      su:
        Language models
        Cognitive testing
        English language
        Chinese language
      sug:
        subj:
          Language models
          Cognitive testing
          English language
          Chinese language
      keyword:
        Brain encoding
        Communication and Culture Linguistics
        Eye-tracking
        FMRI
        Language model
        MEG
        Psychology and Cognitive Sciences Psychology Cognitive Sciences Language
        Semantic rating
      ab: Pre-trained computational language models have recently made remarkable progress in harnessing the language abilities which were considered unique to humans. Their success has raised interest in whether these models represent and process language like humans. To answer this question, this paper proposes MulCogBench, a multi-modal cognitive benchmark dataset collected from native Chinese and English participants. It encompasses a variety of cognitive data, including subjective semantic ratings, eye-tracking, functional magnetic resonance imaging (fMRI), and magnetoencephalography (MEG). To assess the relationship between language models and cognitive data, we conducted a similarity-encoding analysis which decodes cognitive data based on its pattern similarity with textual embeddings. We then validated the reliability of the encoding results by applying the ridge regression to perform the same encoding analysis. Results show that language models share significant similarities with human cognitive data and the similarity patterns are modulated by the data modality and stimuli complexity. Specifically, context-aware models outperform context-independent models as language stimulus complexity increases. The shallow layers of context-aware models are better aligned with the high-temporal-resolution MEG signals whereas the deeper layers show more similarity with the high-spatial-resolution fMRI. These results indicate that language models have a delicate relationship with brain language representations. Moreover, the results between Chinese and English are highly consistent, suggesting the generalizability of these findings across languages.
      pubtype: Academic Journal
      doctype: Conference Paper/Materials
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N