Comparison between GPT-4 and human raters in grading pharmacy students' exam responses in Malaysia: a cross-sectional study.
Purpose: Manual grading is time-consuming and prone to inconsistencies, prompting the exploration of generative artificial intelligence tools such as GPT-4 to enhance efficiency and reliability. This study investigated GPT-4's potential in grading pharmacy students' exam responses, focusing on the i...
| Publicado en: | Journal of Educational Evaluation for Health Professions Vol. 22; pp. 1 - 9 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | pictorial research tables/charts Journal Article |
| Publicado: |
National Health Personnel Licensing Examination Board
2025
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=192170955&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 192170955 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 19755937 B0CC jtl: Journal of Educational Evaluation for Health Professions issn: 19755937 maglogo: N pubinfo: dt: 2025 vid: 22 pid: 58691 pub: National Health Personnel Licensing Examination Board artinfo: ui: 192170955 192170955 192170955 10.3352/jeehp.2025.22.20 192170955 ppf: 1 ppct: 8 formats: fmt: @attributes: type: P tig: atl: Comparison between GPT-4 and human raters in grading pharmacy students' exam responses in Malaysia: a cross-sectional study. aug: au: Yap, Wuan Shuen Saw, Pui San Yeap, Li Ling Huey Lee, Shaun Wen Wong, Wei Jin Seng Lee, Ronald Fook affil: School of Pharmacy, Monash University Malaysia, Bandar Sunway, Malaysia sug: subj: Education, Pharmacy Malaysia Students, Pharmacy Educational Measurement Artificial Intelligence, Generative Reproducibility of Results Human Malaysia Cross Sectional Studies Comparative Studies Colleges and Universities Intraclass Correlation Coefficient Wilcoxon Signed Rank Test Statistical Significance Data Analysis Software Descriptive Statistics Funding Source ab: Purpose: Manual grading is time-consuming and prone to inconsistencies, prompting the exploration of generative artificial intelligence tools such as GPT-4 to enhance efficiency and reliability. This study investigated GPT-4's potential in grading pharmacy students' exam responses, focusing on the impact of optimized prompts. Specifically, it evaluated the alignment between GPT-4 and human raters, assessed GPT-4's consistency over time, and determined its error rates in grading pharmacy students' exam responses. Methods: We conducted a comparative study using past exam responses graded by university-trained raters and by GPT-4. Responses were randomized before evaluation by GPT-4, accessed via a Plus account between April and September 2024. Prompt optimization was performed on 16 responses, followed by evaluation of 3 prompt delivery methods. We then applied the optimized approach across 4 item types. Intraclass correlation coefficients and error analyses were used to assess consistency and agreement between GPT-4 and human ratings. Results: GPT-4's ratings aligned reasonably well with human raters, demonstrating moderate to excellent reliability (intraclass correlation coefficient= 0.617-0.933), depending on item type and the optimized prompt. When stratified by grade bands, GPT-4 was less consistent in marking high-scoring responses (Z= -5.71-4.62, P<0.001). Overall, despite achieving substantial alignment with human raters in many cases, discrepancies across item types and a tendency to commit basic errors necessitate continued educator involvement to ensure grading accuracy. Conclusion: With optimized prompts, GPT-4 shows promise as a supportive tool for grading pharmacy students' exam responses, particularly for objective tasks. However, its limitations--including errors and variability in grading high-scoring responses--require ongoing human oversight. Future research should explore advanced generative artificial intelligence models and broader assessment formats to further enhance grading reliability. pubtype: Academic Journal doctype: pictorial research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|