AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory
2026Original research

Large language models for generating key-feature questions in medical education

Yavuz Selim Kıyak, Stanisław Górski, Tomasz Tokarek, Michał Pers, Andrzej A Kononowicz. Large language models for generating key-feature questions in medical education. Medical Education Online. 2025. https://doi.org/10.1080/10872981.2025.2574647

DOI: 10.1080/10872981.2025.2574647

URL: https://doi.org/10.1080/10872981.2025.2574647

Journal: Medical Education Online

Authors: Yavuz Selim Kıyak, Stanisław Górski, Tomasz Tokarek, Michał Pers, Andrzej A Kononowicz

Last updated: June 28, 2026

Article interpretation

Main contribution
This descriptive study evaluated whether OpenAI’s o3 model could generate key-feature questions aligned with Medical Council of Canada guidance. Using a structured prompt, the authors generated 20 cardiology KFQs from recent ESC guidelines; expert review found 93.7% checklist compliance, with 3 accepted as is and 17 accepted with minor revisions.
Practical use
Educators may use a structured LLM prompt to create first drafts of KFQs, especially when building cases across large content areas. The article supports using AI to reduce drafting burden, not to replace clinical and assessment review.
Main caution
The generated KFQs still had clinically important errors: weak or implausible distractors, missing or incorrect “killer” responses, and occasional inaccuracies such as a pregnancy test option for a male patient. Expert review remains necessary before use.
Research gap
The study tested only 20 English-language cardiology KFQs generated with o3. It did not test learner performance, reliability, discrimination, validity impact, other specialties, other languages, or lower-performing models.

Collaboration

Working on a study that could address this gap?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Curator note

This is primary (and only as of July 2026) evidence on AI-generated KFQs.

Caution

Do not treat the 93.7% compliance rate as evidence that LLM-generated KFQs are ready for high-stakes exams. The study assessed expert-rated item quality, not psychometric performance in real assessment settings.

Prompt / method

Includes prompt
Yes

Related links

Related main question areas: What can AI do for generating assessment materials?

Related subquestions: What can AI do for generating key-feature questions?

Related published prompts: None linked

Tags: key-feature questions