AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory

What can AI do for generating MCQs?

AI can produce rapid first drafts of MCQs when given input such as source content, learning objectives, item format. AI should not be treated as an autonomous item writer. The current evidence supports human-in-the-loop item development, not unsupervised use in summative or high-stakes exams.

Evidence summary

A systematic review of 71 medical education studies found that AI-generated MCQs are feasible, fast, and often acceptable after review, but validity evidence remains incomplete. AI-generated items often show moderate difficulty, but discrimination is modest. Non-functional distractors are common, meaning some answer options are rarely chosen and do not test effectively. Reported factual or content error rates varied widely, from less than 1% to 45%. Nearly half of studies relied only on expert review, without testing items on learners. AI can produce major time and cost savings, but the savings mainly occur in drafting.

Practical starting points

  • Start with a blueprint, learning objective, source content, target learner level, and item-writing rules.
  • Ask AI to generate structured draft items, including correct answer and rationale.
  • Have subject experts check factual accuracy, relevance, cognitive level, and answer defensibility.
  • Review stems and distractors for ambiguity, cueing, implausibility, cultural mismatch, and option-pattern bias.
  • Pilot items with learners where possible.
  • Analyze difficulty, discrimination, distractor function, and reliability.
  • Revise, approve, retain, or remove items based on evidence.

Last updated: July 2, 2026

Collaboration

Planning a study related to this question?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this question or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Related articles

Selected studies

2026Systematic review

Validity of AI-generated multiple-choice questions in medical education: a systematic review

Yavuz Selim Kıyak, Abdullah Bedir Kaya, Emre Emekli · Postgraduate Medical Journal

This review organizes the evidence on AI-generated MCQs in medical education using Messick’s validity framework. It shows that evidence is strongest for content review, but weaker for response processes, psychometrics, links with other assessments, and consequences.

Related subquestions: What can AI do for generating MCQs?

Published prompts

Related prompt architectures

Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License

Medical MCQ generation prompt template (case-based)

Prompt

Notes: Original prompt for generating case-based single-best-answer MCQs in medical education. Adapted into a Custom GPT while preserving the original instructional framework. Intended as a prompt resource rather than evidence of effectiveness.

Related subquestions: What can AI do for generating MCQs?

Open prompt page
Fair use

NBME-style clinical vignette MCQ prompt template

Prompt

Notes: Prompt for generating NBME-style single-best-answer multiple-choice questions using authentic clinical vignettes. Asks the model to include a four-sentence patient scenario, vital signs, and physical examination findings while avoiding pseudovignettes. Requires an item-specific test point to define the learning objective.

Related subquestions: What can AI do for generating MCQs?

Open prompt page

Practical guides

Linked guides

No guides are linked yet.