Comparison between GPT-4 and human raters in grading pharmacy students' exam responses in Malaysia: a cross-sectional study.

Purpose: Manual grading is time-consuming and prone to inconsistencies, prompting the exploration of generative artificial intelligence tools such as GPT-4 to enhance efficiency and reliability. This study investigated GPT-4's potential in grading pharmacy students' exam responses, focusing on the i...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Educational Evaluation for Health Professions Vol. 22; pp. 1 - 9
Autores principales: Yap, Wuan Shuen, Saw, Pui San, Yeap, Li Ling, Huey Lee, Shaun Wen, Wong, Wei Jin, Seng Lee, Ronald Fook
Formato: pictorial research tables/charts Journal Article
Publicado: National Health Personnel Licensing Examination Board 2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=192170955&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 192170955
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        19755937
        B0CC
      jtl: Journal of Educational Evaluation for Health Professions
      issn: 19755937
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 22
      pid: 58691
      pub: National Health Personnel Licensing Examination Board
    artinfo:
      ui:
        192170955
        192170955
        192170955
        10.3352/jeehp.2025.22.20
        192170955
      ppf: 1
      ppct: 8
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Comparison between GPT-4 and human raters in grading pharmacy students' exam responses in Malaysia: a cross-sectional study.
      aug:
        au:
          Yap, Wuan Shuen
          Saw, Pui San
          Yeap, Li Ling
          Huey Lee, Shaun Wen
          Wong, Wei Jin
          Seng Lee, Ronald Fook
        affil: School of Pharmacy, Monash University Malaysia, Bandar Sunway, Malaysia
      sug:
        subj:
          Education, Pharmacy Malaysia
          Students, Pharmacy
          Educational Measurement
          Artificial Intelligence, Generative
          Reproducibility of Results
          Human
          Malaysia
          Cross Sectional Studies
          Comparative Studies
          Colleges and Universities
          Intraclass Correlation Coefficient
          Wilcoxon Signed Rank Test
          Statistical Significance
          Data Analysis Software
          Descriptive Statistics
          Funding Source
      ab: Purpose: Manual grading is time-consuming and prone to inconsistencies, prompting the exploration of generative artificial intelligence tools such as GPT-4 to enhance efficiency and reliability. This study investigated GPT-4's potential in grading pharmacy students' exam responses, focusing on the impact of optimized prompts. Specifically, it evaluated the alignment between GPT-4 and human raters, assessed GPT-4's consistency over time, and determined its error rates in grading pharmacy students' exam responses. Methods: We conducted a comparative study using past exam responses graded by university-trained raters and by GPT-4. Responses were randomized before evaluation by GPT-4, accessed via a Plus account between April and September 2024. Prompt optimization was performed on 16 responses, followed by evaluation of 3 prompt delivery methods. We then applied the optimized approach across 4 item types. Intraclass correlation coefficients and error analyses were used to assess consistency and agreement between GPT-4 and human ratings. Results: GPT-4's ratings aligned reasonably well with human raters, demonstrating moderate to excellent reliability (intraclass correlation coefficient= 0.617-0.933), depending on item type and the optimized prompt. When stratified by grade bands, GPT-4 was less consistent in marking high-scoring responses (Z= -5.71-4.62, P<0.001). Overall, despite achieving substantial alignment with human raters in many cases, discrepancies across item types and a tendency to commit basic errors necessitate continued educator involvement to ensure grading accuracy. Conclusion: With optimized prompts, GPT-4 shows promise as a supportive tool for grading pharmacy students' exam responses, particularly for objective tasks. However, its limitations--including errors and variability in grading high-scoring responses--require ongoing human oversight. Future research should explore advanced generative artificial intelligence models and broader assessment formats to further enhance grading reliability.
      pubtype: Academic Journal
      doctype:
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N