AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory
2026Review articleSystematic review

Validity of AI-generated multiple-choice questions in medical education: a systematic review

Yavuz Selim Kıyak, Abdullah Bedir Kaya, Emre Emekli. Validity of AI-generated multiple-choice questions in medical education: a systematic review. Postgraduate Medical Journal. 2026. https://doi.org/10.1093/POSTMJ/QGAG057

Abstract or concise summary

Large language models (LLMs) are increasingly used to generate multiple-choice questions (MCQs) in medical education. We conducted a systematic review following PRISMA 2020, searching PubMed, Web of Science, Scopus, and ERIC through 15 February 2026. Two reviewers independently screened 1352 records, extracted data and assessed methodological quality using the Joanna Briggs Institute checklist. Findings were synthesized according to Messick’s five sources of validity evidence. Seventy-one studies from 24 countries were included. Most relied on expert review rather than learner-based testing. All studies reported content evidence, whereas fewer addressed relations to other variables (40/71), response process (35/71), internal structure (31/71) or consequences (25/71). Error rates ranged from <1% to 45%. Median item difficulty was 0.67 and discrimination 0.28, with reliability between 0.51–0.81. Studies reported substantial efficiency gains, including up to 31-fold time savings. LLMs appear useful drafting tools but current evidence does not yet support unsupervised use in summative assessment.

DOI: 10.1093/POSTMJ/QGAG057

URL: https://doi.org/10.1093/postmj/qgag057

Journal: Postgraduate Medical Journal

Authors: Yavuz Selim Kıyak, Abdullah Bedir Kaya, Emre Emekli

Last updated: July 5, 2026

Article interpretation

Main contribution
This review organizes the evidence on AI-generated MCQs in medical education using Messick’s validity framework. It shows that evidence is strongest for content review, but weaker for response processes, psychometrics, links with other assessments, and consequences.
Practical use
Read the article for understanding the latest evidence on using AI for generating MCQs.
Main caution
AI-generated MCQs can look polished while still containing factual errors, weak distractors, cueing, poor discrimination, or misalignment with the intended construct. Time savings mainly occur at the drafting stage; expert review and governance remain necessary. Current evidence does not support unsupervised use in high-stakes or summative exams.
Research gap
The field needs stronger learner-based evidence, standardized quality-rating tools, transparent reporting of model versions and prompts, and studies on fairness, security, learning effects, and long-term assessment consequences.

Collaboration

Working on a study that could address this gap?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Curator note

This study includes the most up-to-date findings as of July 2026. Its search cut-off date is 15 February 2026.

Caution

Most included studies evaluated older OpenAI models (GPT-3.5, GPT-4, and GPT-4o) while open-source and newer models were rarely examined.

Related links