Prompt engineering for single-best-answer multiple-choice questions in licensing examinations: a narrative review with a case study involving the Korean Medical Licensing Examination.
The emergence of large language models (LLMs) has generated growing interest in their potential applications for medical assessment and item development. This practice-oriented narrative review examines the potential of LLMs, particularly ChatGPT, for generating and validating single-best-answer mul...
| Publicado en: | Journal of Educational Evaluation for Health Professions Vol. 22; pp. 1 - 16 |
|---|---|
| Autores principales: | , , , |
| Formato: | case study review Journal Article |
| Publicado: |
National Health Personnel Licensing Examination Board
2025
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=192170969&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 192170969 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 19755937 B0CC jtl: Journal of Educational Evaluation for Health Professions issn: 19755937 maglogo: N pubinfo: dt: 2025 vid: 22 pid: 58691 pub: National Health Personnel Licensing Examination Board artinfo: ui: 192170969 192170969 192170969 10.3352/jeehp.2025.22.34 192170969 ppf: 1 ppct: 15 formats: fmt: @attributes: type: P tig: atl: Prompt engineering for single-best-answer multiple-choice questions in licensing examinations: a narrative review with a case study involving the Korean Medical Licensing Examination. aug: au: Kim, Bokyoung Kang, Junseok Kim, Min-Young Ahn, Jihyun affil: College of Nursing, Research Institute of Nursing Innovation, Kyungpook National University, Daegu, Korea sug: subj: Licensure Artificial Intelligence Educational Measurement Test Construction Natural Language Processing Information Retrieval Checklists Quality Assurance Reproducibility of Results Educational Technology ab: The emergence of large language models (LLMs) has generated growing interest in their potential applications for medical assessment and item development. This practice-oriented narrative review examines the potential of LLMs, particularly ChatGPT, for generating and validating single-best-answer multiple-choice questions in health professions licensing examinations, using a Korean Medical Licensing Examination (KMLE)-focused case perspective. We frame LLMs as human-in-the-loop tools rather than replacements for high-stakes testing. Recent applications of LLMs in assessment were reviewed, including prompting strategies such as few-shot, multi-stage, and chain-of-thought methods, as well as retrieval-augmented generation (RAG) to align outputs with exam blueprints. Approaches to enforcing formatting rules, checklist-based self-validation, and iterative refinement were analyzed for their role in supporting item development. Findings indicate that LLMs can perform near passing thresholds on high-stakes exams and assist with grading and feedback tasks. Prompt engineering enhances structural fidelity and clinical plausibility, while human oversight remains critical for accuracy, cultural appropriateness, and psychometric defensibility. The emerging multimodal generation of images, audio, and video suggests the feasibility of new item formats, provided robust validation safeguards are implemented. The most effective approach is a human-in-the-loop workflow that leverages artificial intelligence efficiency while embedding expert judgment, psychometric evaluation, and ethical governance. This practice-oriented roadmap--integrating strategic prompt selection, RAG-based blueprint alignment, rigorous validation gates, and KMLE-specific formatting--offers an implementable and methodologically defensible approach for licensing examinations. pubtype: Academic Journal doctype: case study review Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|