Validity of AI-generated multiple-choice questions in medical education: a systematic review
Yavuz Selim Kıyak, Abdullah Bedir Kaya, Emre Emekli. Validity of AI-generated multiple-choice questions in medical education: a systematic review. Postgraduate Medical Journal. 2026. https://doi.org/10.1093/POSTMJ/QGAG057
Abstract or concise summary
Large language models (LLMs) are increasingly used to generate multiple-choice questions (MCQs) in medical education. We conducted a systematic review following PRISMA 2020, searching PubMed, Web of Science, Scopus, and ERIC through 15 February 2026. Two reviewers independently screened 1352 records, extracted data and assessed methodological quality using the Joanna Briggs Institute checklist. Findings were synthesized according to Messick’s five sources of validity evidence. Seventy-one studies from 24 countries were included. Most relied on expert review rather than learner-based testing. All studies reported content evidence, whereas fewer addressed relations to other variables (40/71), response process (35/71), internal structure (31/71) or consequences (25/71). Error rates ranged from <1% to 45%. Median item difficulty was 0.67 and discrimination 0.28, with reliability between 0.51–0.81. Studies reported substantial efficiency gains, including up to 31-fold time savings. LLMs appear useful drafting tools but current evidence does not yet support unsupervised use in summative assessment.
DOI: 10.1093/POSTMJ/QGAG057
URL: https://doi.org/10.1093/postmj/qgag057
Journal: Postgraduate Medical Journal
Authors: Yavuz Selim Kıyak, Abdullah Bedir Kaya, Emre Emekli
Last updated: July 5, 2026
Article interpretation
- Main contribution
- This review organizes the evidence on AI-generated MCQs in medical education using Messick’s validity framework. It shows that evidence is strongest for content review, but weaker for response processes, psychometrics, links with other assessments, and consequences.
- Practical use
- Read the article for understanding the latest evidence on using AI for generating MCQs.
- Main caution
- AI-generated MCQs can look polished while still containing factual errors, weak distractors, cueing, poor discrimination, or misalignment with the intended construct. Time savings mainly occur at the drafting stage; expert review and governance remain necessary. Current evidence does not support unsupervised use in high-stakes or summative exams.
- Research gap
- The field needs stronger learner-based evidence, standardized quality-rating tools, transparent reporting of model versions and prompts, and studies on fairness, security, learning effects, and long-term assessment consequences.
Collaboration
Working on a study that could address this gap?
If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.
Curator note
This study includes the most up-to-date findings as of July 2026. Its search cut-off date is 15 February 2026.
Caution
Related links
Related main question areas: What can AI do for generating assessment materials?
Related subquestions: What can AI do for generating MCQs?
Related published prompts: Medical MCQ generation prompt template (case-based), NBME-style clinical vignette MCQ prompt template
Tags: systematic review, MCQ, validity