AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory

AI MedEd Observatory

What Can AI Do in Medical Education?

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

A curated resource for medical educators and researchers who want to understand how AI is being used and what the evidence suggests.

Curated by Yavuz Selim Kıyak, MD, PhD

As a researcher deeply involved in integrating AI into medical education, with several publications in Medical Teacher, Medical Education, and Academic Medicine, I created this personally curated resource because the literature on AI in medical education is expanding rapidly, making it increasingly difficult to identify what is useful, trustworthy, and applicable. This site organizes selected evidence and shares what I am learning. It is to help educators and researchers.

Mantra

Evidence before automation. Judgment before adoption.

Main question areas

Practical starting points

The observatory is organized around common educational questions rather than around AI tools.

4 subquestions5 linked articles5 published prompts

What can AI do for generating assessment materials?

Explore area
1 subquestions2 linked articles0 published prompts

What can AI do for automated scoring?

Explore area
1 subquestions2 linked articles0 published prompts

What can AI do for feedback?

Explore area

Why this site exists

Practical interpretation, not a generic tool directory

Many articles now report AI-generated questions, feedback, scoring, and educational materials. This site organizes selected studies into practical, evidence-informed summaries so educators and researchers can judge what may be useful, what requires expert review, and what remains uncertain.

Curator note

Human review remains essential. Use as a starting point.

Caution

Current evidence suggests some uses may help with drafting, structuring, or calibration. Evidence is limited, and many applications are not ready for unsupervised summative use.

Trust this resource?

Context for curation

This Observatory is curated by Yavuz Selim Kıyak, MD, PhD. His work in this area was noted in an independent bibliometric analysis of ChatGPT in medical education, which identified him as the most prolific author in the field. This is shared only as context for the curation of the resource, not as a claim of authority. The value of this Observatory lies in its transparent organization of the literature, practical interpretation of published evidence, and direct links to the original studies, allowing readers to form their own informed judgments.

Latest additions

Recently added articles and prompts

Recent entries help show where the observatory is currently growing.

Article

Added July 19, 2026

Large language models in professionalism and ethical decision-making education: a transferable process for generating concordance of professional judgment items

This innovation report presents an evaluated human-in-the-loop workflow for generating Concordance of Professional Judgment items from a locally defined ethics framework. ChatGPT-5 Thinking and Gemini 2.5 Pro each generated 28 Turkish items; after independent faculty review and adjudication, 25/28 and 27/28 items, respectively, were accepted or required only minor revision.

Read article summary

Article

Added July 5, 2026

Is AI the future of evaluation in medical education?? AI vs. human evaluation in objective structured clinical examination

This cross-sectional study compared OSCE scores from three human evaluators and two AI systems, ChatGPT-4o and Gemini Flash 1.5, for 196 pre-clinical medical students performing four clinical skills. AI scores were generally higher than human scores, with better alignment on visually observable steps than on auditory or communication-dependent criteria.

Read article summary

Article

Added July 5, 2026

Validity of AI-generated multiple-choice questions in medical education: a systematic review

This review organizes the evidence on AI-generated MCQs in medical education using Messick’s validity framework. It shows that evidence is strongest for content review, but weaker for response processes, psychometrics, links with other assessments, and consequences.

Read article summary
Creative Commons Attribution 4.0 International License (CC BY 4.0)

Concordance of Professional Judgment item generation prompt template

Added July 19, 2026

Prompt

Notes: Original full prompt evaluated with ChatGPT-5 Thinking and Gemini 2.5 Pro to generate CoPJ item drafts across 28 topics in the Turkish Medical Association's ethical declarations. Intended for formative content development with local parameter adaptation and expert review.

Related subquestions: What can AI do for generating Concordance of Professional Judgment items?

Open prompt page
Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License

Medical MCQ generation prompt template (case-based)

Added July 5, 2026

Prompt

Notes: Original prompt for generating case-based single-best-answer MCQs in medical education. Adapted into a Custom GPT while preserving the original instructional framework. Intended as a prompt resource rather than evidence of effectiveness.

Related subquestions: What can AI do for generating MCQs?

Open prompt page