Assessing the efficacy of LLMs in neuroscience: A performance analysis of ChatGPT and gemini on synaptic plasticity.

Synaptic plasticity, which plays a critical role in fundamental neurological processes, is a complex subject to master. Therefore, large language models (LLMs) are increasingly being used to facilitate the learning of such complex topics. However, these models have limitations, including producing i...

Descripción completa

Detalles Bibliográficos
Publicado en:Work p. 1
Autores principales: Altunkaya, Melek, Babur, Ercan, Cihan, Emine, Sahbaz Pirincci, Cansu
Formato: Journal Article
Publicado: Sage Publications Inc. Jul2026
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Synaptic plasticity, which plays a critical role in fundamental neurological processes, is a complex subject to master. Therefore, large language models (LLMs) are increasingly being used to facilitate the learning of such complex topics. However, these models have limitations, including producing inaccurate information and failing to capture the nuances of scientific terminology.This study aimed to evaluate the accuracy, quality and readability of LLM responses to questions on synaptic plasticity.The widely used LLMs ChatGPT-4 and Gemini 2.5 were selected in the study. Ten questions were posed to each LLM, and the initial responses were recorded. Five neurophysiologists evaluated the responses qualitatively using a 4-point Likert scale. Readability level of the answers was analyzed using Flesch-Kincaid Grade Level test.In the qualitative assessment, both models generally provided accurate and acceptable information. Within the limited scope of the questions analyzed, Gemini received higher median scores in certain instances; however, no statistically significant difference was observed between the two models across most of the question set. Linguistic analysis showed that Gemini's responses were longer and featured a higher Flesch-Kincaid Grade Level, suggesting a structure more aligned with academic or technical discourse.For the specific neuroscientific inquiries examined in this study, both LLMs demonstrated a high capacity for generating accurate content. While Gemini's responses exhibited a more technical linguistic profile, the findings are context-specific and further research is needed to determine if these trends persist across broader scientific domains and larger datasets.