AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory

Evidence map

Search the article record

Article cards summarize titles, study type, whether prompts are included, practical uses, cautions, and related subquestions.

8 articlesSearchable by title, author, journal, tag, abstract, contribution, and curator note
2026Other

Large language models in professionalism and ethical decision-making education: a transferable process for generating concordance of professional judgment items

Tuğba İş-Kara, Yavuz Selim Kıyak · Academic Medicine

This innovation report presents an evaluated human-in-the-loop workflow for generating Concordance of Professional Judgment items from a locally defined ethics framework. ChatGPT-5 Thinking and Gemini 2.5 Pro each generated 28 Turkish items; after independent faculty review and adjudication, 25/28 and 27/28 items, respectively, were accepted or required only minor revision.

Related subquestions: What can AI do for generating Concordance of Professional Judgment items?

2026Other

Automated scoring of student videos in medical education: a comparison between a large language model and expert evaluation

Yavuz Selim Kıyak, Özlem Ülkü Bulut, Özlem Coşkun, Işıl İrem Budakoğlu · Journal of Microbiology & Biology Education

This study shows that Gemini 2.5 Pro did not reliably match expert scoring of 139 medical student EBM video presentations. A rubric-only prompt over-scored students, while a “critical expert” prompt under-scored them, showing that prompt strategy changed both the size and direction of scoring bias.

Related subquestions: What can AI do for scoring videos?

2026Original research

Large language models for generating key-feature questions in medical education

Yavuz Selim Kıyak, Stanisław Górski, Tomasz Tokarek, Michał Pers, Andrzej A Kononowicz · Medical Education Online

This descriptive study evaluated whether OpenAI’s o3 model could generate key-feature questions aligned with Medical Council of Canada guidance. Using a structured prompt, the authors generated 20 cardiology KFQs from recent ESC guidelines; expert review found 93.7% checklist compliance, with 3 accepted as is and 17 accepted with minor revisions.

Related subquestions: What can AI do for generating key-feature questions?

2026Scoping review

Applications and Outcomes of Large‑Language‑Model‑Generated Feedback in Undergraduate Medical Education: A Scoping Review

Yavuz Selim Kıyak, Tuğba İş-Kara, Emre Emekli · Medical Science Educator

This scoping review maps 42 peer-reviewed studies on LLM-generated feedback in undergraduate medical education, showing that LLMs are being used across MCQs, clinical reasoning, simulated patients, communication skills, academic writing, and self-directed learning. Evidence mainly supports feasibility and short-term learning or learner reaction, not behavior change or patient outcomes.

Related subquestions: What can AI do for generating feedback?

2026Systematic review

Validity of AI-generated multiple-choice questions in medical education: a systematic review

Yavuz Selim Kıyak, Abdullah Bedir Kaya, Emre Emekli · Postgraduate Medical Journal

This review organizes the evidence on AI-generated MCQs in medical education using Messick’s validity framework. It shows that evidence is strongest for content review, but weaker for response processes, psychometrics, links with other assessments, and consequences.

Related subquestions: What can AI do for generating MCQs?

2025Original research

Is AI the future of evaluation in medical education?? AI vs. human evaluation in objective structured clinical examination

Murat Tekin, Mustafa Onur Yurdal, Çetin Toraman, Güneş Korkmaz, İbrahim Uysal · BMC Medical Education

This cross-sectional study compared OSCE scores from three human evaluators and two AI systems, ChatGPT-4o and Gemini Flash 1.5, for 196 pre-clinical medical students performing four clinical skills. AI scores were generally higher than human scores, with better alignment on visually observable steps than on auditory or communication-dependent criteria.

Related subquestions: What can AI do for scoring videos?

2025

Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude

Zuhal Yapıcı Coşkun, Yavuz Selim Kıyak, Özlem Coşkun, Işıl İrem Budakoğlu, Özhan Özdemir · Medical Teacher

This cross-sectional study evaluated ChatGPT-4o and Claude 3.5 Sonnet for generating obstetrics and gynecology SCT items on five primary-care diagnostic topics. Sixteen obstetrics and gynecology residents rated 10 AI-generated SCT items against 11 criteria; overall agreement that items met criteria was 90.57% for ChatGPT-4o and 91.48% for Claude 3.5 Sonnet.

Related subquestions: What can AI do for generating script concordance tests?

2024Other

Using Large Language Models to Generate Script Concordance Test in Medical Education: ChatGPT and Claude

Yavuz Selim Kıyak, Emre Emekli · Revista Española de Educación Médica

This study tested whether ChatGPT-4 and Claude 3 Sonnet could generate Script Concordance Test items for abdominal radiology using a detailed prompt. Sixteen radiologists judged the items; most rated them as assessing clinical reasoning rather than factual recall, and ChatGPT-4 had higher acceptability than Claude across the reported indicators.

Related subquestions: What can AI do for generating script concordance tests?