AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory

What can AI do for generating key-feature questions?

AI can produce structured first drafts of key-feature questions when given a clinical problem, candidate level, patient age group, clinical situation, site of care, and a reference guideline. AI should not be treated as an autonomous KFQ writer. The current evidence supports human-in-the-loop KFQ development, not unsupervised use in summative or high-stakes exams.

Evidence summary

A descriptive study evaluated 20 AI-generated cardiology KFQs created with OpenAI’s o3 model using a structured prompt aligned with Medical Council of Canada KFQ guidelines. Each KFQ was reviewed by two cardiology experts, with disagreements resolved by a third reviewer. Of the 20 KFQs, 3 were accepted as is and 17 were accepted with minor revisions. None required major revisions or rejection. Overall checklist compliance was 93.7%. The strongest areas were key-feature definition, clinical scenario plausibility, alignment between scenario and questions, and focus on clinical decision-making rather than simple recall. The main weaknesses were clinically harmful “killer” responses, plausible short-menu distractors, and active decision-making language in question wording. Some items included implausible distractors, missing clinical variables, contradictions between questions, or incorrect classification of harmful options. Psychometric performance with learners was not tested.

Practical starting points

  • Start with a clinical problem, candidate level, patient age group, clinical situation, site of care, and reference guideline.
  • Ask AI to generate the full KFQ structure using the prompt template.
  • Review questions by using Kiyak’s quality checklist for evaluating Key-Feature Questions.

Last updated: July 3, 2026

Collaboration

Planning a study related to this question?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this question or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Related articles

Selected studies

2026Original research

Large language models for generating key-feature questions in medical education

Yavuz Selim Kıyak, Stanisław Górski, Tomasz Tokarek, Michał Pers, Andrzej A Kononowicz · Medical Education Online

This descriptive study evaluated whether OpenAI’s o3 model could generate key-feature questions aligned with Medical Council of Canada guidance. Using a structured prompt, the authors generated 20 cardiology KFQs from recent ESC guidelines; expert review found 93.7% checklist compliance, with 3 accepted as is and 17 accepted with minor revisions.

Related subquestions: What can AI do for generating key-feature questions?

Published prompts

Related prompt architectures

Creative Commons Attribution-NonCommercial License 4.0

Prompt for generating key-feature questions (problems) in medical education

Prompt

Related subquestions: What can AI do for generating key-feature questions?

Open prompt page

Practical guides

Linked guides