Is ChatGPT 'ready' to be a learning tool for medical undergraduates and will it perform equally in different subjects? Comparative study of ChatGPT performance in tutorial and case-based learning questions in physiology and biochemistry.

Purpose: Generative AI will become an integral part of education in future. The potential of this technology in different disciplines should be identified to promote effective adoption. This study evaluated the performance of ChatGPT in tutorial and case-based learning questions in physiology and bi...

Descripción completa

Detalles Bibliográficos
Publicado en:Medical Teacher Vol. 46; no. 11; pp. 1441 - 1448
Autores principales: Luke, W. A. Nathasha V., Seow Chong, Lee, Ban, Kenneth H., Wong, Amanda H., Zhi Xiong, Chen, Shuh Shing, Lee, Taneja, Reshma, Samarasekera, Dujeepa D., Yap, Celestial T.
Formato: research tables/charts Journal Article
Publicado: Taylor & Francis Ltd Nov2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=180625209&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 180625209
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        0142159X
        MCH
      jtl: Medical Teacher
      issn: 0142159X
      maglogo: Y
    pubinfo:
      dt: Nov2024
      vid: 46
      iid: 11
      pid: 377
      pub: Taylor & Francis Ltd
      place: Philadelphia, Pennsylvania
    artinfo:
      ui:
        180625209
        175154543
        180625209
        180625209
        10.1080/0142159X.2024.2308779
        180625209
      ppf: 1441
      ppct: 7
      formats:
      tig:
        atl: Is ChatGPT 'ready' to be a learning tool for medical undergraduates and will it perform equally in different subjects? Comparative study of ChatGPT performance in tutorial and case-based learning questions in physiology and biochemistry.
      aug:
        au:
          Luke, W. A. Nathasha V.
          Seow Chong, Lee
          Ban, Kenneth H.
          Wong, Amanda H.
          Zhi Xiong, Chen
          Shuh Shing, Lee
          Taneja, Reshma
          Samarasekera, Dujeepa D.
          Yap, Celestial T.
        affil: Department of Physiology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore
      sug:
        subj:
          Chatbot
          Students, Medical
          Students, Undergraduate
          Learning Methods
          Biochemistry Education
          Problem-Based Learning
          Teaching Methods, Clinical
          Education, Medical
          Physiology Education
          Human
          Student Performance Appraisal
          Natural Language Processing
          Faculty, Medical
          Bloom's Taxonomy
          Funding Source
      ab: Purpose: Generative AI will become an integral part of education in future. The potential of this technology in different disciplines should be identified to promote effective adoption. This study evaluated the performance of ChatGPT in tutorial and case-based learning questions in physiology and biochemistry for medical undergraduates. Our study mainly focused on the performance of GPT-3.5 version while a subgroup was comparatively assessed on GPT-3.5 and GPT-4 performances. Materials and methods: Answers were generated in GPT-3.5 for 44 modified essay questions (MEQs) in physiology and 43 MEQs in biochemistry. Each answer was graded by two independent examiners. Subsequently, a subset of 15 questions from each subject were selected to represent different score categories of the GPT-3.5 answers; responses were generated in GPT-4, and graded. Results: The mean score for physiology answers was 74.7 (SD 25.96). GPT-3.5 demonstrated a statistically significant (p =.009) superior performance in lower-order questions of Bloom's taxonomy in comparison to higher-order questions. Deficiencies in the application of physiological principles in clinical context were noted as a drawback. Scores in biochemistry were relatively lower with a mean score of 59.3 (SD 26.9) for GPT-3.5. There was no statistically significant difference in the scores for higher and lower-order questions of Bloom's taxonomy. The deficiencies highlighted were lack of in-depth explanations and precision. The subset of questions where the GPT-4 and GPT-3.5 were compared demonstrated a better overall performance in GPT-4 responses in both subjects. This difference between the GPT-3.5 and GPT-4 performance was statistically significant in biochemistry but not in physiology. Conclusions: The differences in performance across the two versions, GPT-3.5 and GPT-4 across the disciplines are noteworthy. Educators and students should understand the strengths and limitations of this technology in different fields to effectively integrate this technology into teaching and learning.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N