Comparative Evaluation of ChatGPT-4o, Gemini 2.5 Pro and Grok-4 in Answering Orthodontics Questions from the Dentistry Specialty Examination.
Introduction: This study aimed to compare the accuracy of three advanced Large Language Models (LLMs) in answering orthodontics-related questions from the Turkish Dentistry Specialty Examination (DUS) and to assess their performance across different examination periods. Methods: A total of 129 ortho...
| Publicado en: | Lokman Hekim Health Sciences Vol. 6; no. 1; pp. 126 - 134 |
|---|---|
| Autores principales: | , |
| Formato: | research tables/charts Journal Article |
| Publicado: |
KARE Publishing
Mar2026
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=193195367&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 193195367 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 27917835 N1MV jtl: Lokman Hekim Health Sciences issn: 27917835 maglogo: N pubinfo: dt: Mar2026 vid: 6 iid: 1 pid: 62027 pub: KARE Publishing artinfo: ui: 193195367 193195367 193195367 10.14744/lhhs.2025.42941 193195367 ppf: 126 ppct: 8 formats: fmt: @attributes: type: P tig: atl: Comparative Evaluation of ChatGPT-4o, Gemini 2.5 Pro and Grok-4 in Answering Orthodontics Questions from the Dentistry Specialty Examination. aug: au: Özden, Samet Erener, Hande affil: Department of Orthodontics, İnönü University Faculty of Dentistry, Malatya, Türkiye sug: subj: Prediction Models Artificial Intelligence, Generative Evaluation Natural Language Processing Orthodontics Education Test Taking Dentistry Specialties, Dental Education Credentialing Examinations Turkiye Academic Performance Evaluation Problem Solving Validity Educational Measurement Human Turkiye Education, Dental Descriptive Research Cross Sectional Studies Comparative Studies Computer Simulation Quantitative Studies Descriptive Statistics Data Analysis Software Fisher's Exact Test Pearson's Correlation Coefficient Chi Square Test Post Hoc Analysis Competency Assessment Health Knowledge Evaluation Computer-Assisted Instruction Machine Learning Algorithms ab: Introduction: This study aimed to compare the accuracy of three advanced Large Language Models (LLMs) in answering orthodontics-related questions from the Turkish Dentistry Specialty Examination (DUS) and to assess their performance across different examination periods. Methods: A total of 129 orthodontic questions that were publicly available from 13 DUS sessions conducted between 2012 and 2021 were included. All questions were presented in their original Turkish format, simultaneously, and under identical default settings (i.e., without fine-tuning or additional prompt engineering) by the same operator to eliminate procedural variability. Each model's responses were recorded and scored as correct (1) or incorrect (0). Accuracy comparisons among LLMs were performed using Chi-square and Fisher's exact tests with Monte Carlo correction. Statistical significance was set at p<0.05. Results: No statistically significant differences were observed among the three LLMs within individual examination periods (p>0.05). Grok-4 achieved the highest cumulative accuracy (112/129; 86.8%), followed by Gemini (107/129; 82.9%) and ChatGPT-4o (101/129; 78.3%). The 2018 DUS yielded the lowest accuracy for all models (30%, 30%, and 50%, respectively). All three LLMs performed significantly better on text-based than on figure-based questions (p<0.05), with figure-based accuracy dropping to 45.5% for ChatGPT and Gemini, and 63.6% for Grok. No significant inter-model differences were found within each question type (p>0.05). Discussion and Conclusion: All three LLMs demonstrated high but not flawless accuracy in orthodontics-related DUS questions, with consistent challenges in visual question interpretation. While their integration into examination preparation and dental education holds promise, further refinement in visual reasoning and domain-specific adaptation is needed before clinical or high-stakes implementation. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|