8 articlesSearchable by title, author, journal, tag, abstract, contribution, and curator note
2026Other
Tuğba İş-Kara, Yavuz Selim Kıyak · Academic Medicine
This innovation report presents an evaluated human-in-the-loop workflow for generating Concordance of Professional Judgment items from a locally defined ethics framework. ChatGPT-5 Thinking and Gemini 2.5 Pro each generated 28 Turkish items; after independent faculty review and adjudication, 25/28 and 27/28 items, respectively, were accepted or required only minor revision.
Related subquestions: What can AI do for generating Concordance of Professional Judgment items?
2026Other
Yavuz Selim Kıyak, Özlem Ülkü Bulut, Özlem Coşkun, Işıl İrem Budakoğlu · Journal of Microbiology & Biology Education
This study shows that Gemini 2.5 Pro did not reliably match expert scoring of 139 medical student EBM video presentations. A rubric-only prompt over-scored students, while a “critical expert” prompt under-scored them, showing that prompt strategy changed both the size and direction of scoring bias.
Related subquestions: What can AI do for scoring videos?
2026Original research
Yavuz Selim Kıyak, Stanisław Górski, Tomasz Tokarek, Michał Pers, Andrzej A Kononowicz · Medical Education Online
This descriptive study evaluated whether OpenAI’s o3 model could generate key-feature questions aligned with Medical Council of Canada guidance. Using a structured prompt, the authors generated 20 cardiology KFQs from recent ESC guidelines; expert review found 93.7% checklist compliance, with 3 accepted as is and 17 accepted with minor revisions.
Related subquestions: What can AI do for generating key-feature questions?
2026Scoping review
Yavuz Selim Kıyak, Tuğba İş-Kara, Emre Emekli · Medical Science Educator
This scoping review maps 42 peer-reviewed studies on LLM-generated feedback in undergraduate medical education, showing that LLMs are being used across MCQs, clinical reasoning, simulated patients, communication skills, academic writing, and self-directed learning. Evidence mainly supports feasibility and short-term learning or learner reaction, not behavior change or patient outcomes.
Related subquestions: What can AI do for generating feedback?
2026Systematic review
Yavuz Selim Kıyak, Abdullah Bedir Kaya, Emre Emekli · Postgraduate Medical Journal
This review organizes the evidence on AI-generated MCQs in medical education using Messick’s validity framework. It shows that evidence is strongest for content review, but weaker for response processes, psychometrics, links with other assessments, and consequences.
Related subquestions: What can AI do for generating MCQs?
2025Original research
Murat Tekin, Mustafa Onur Yurdal, Çetin Toraman, Güneş Korkmaz, İbrahim Uysal · BMC Medical Education
This cross-sectional study compared OSCE scores from three human evaluators and two AI systems, ChatGPT-4o and Gemini Flash 1.5, for 196 pre-clinical medical students performing four clinical skills. AI scores were generally higher than human scores, with better alignment on visually observable steps than on auditory or communication-dependent criteria.
Related subquestions: What can AI do for scoring videos?
2025
Zuhal Yapıcı Coşkun, Yavuz Selim Kıyak, Özlem Coşkun, Işıl İrem Budakoğlu, Özhan Özdemir · Medical Teacher
This cross-sectional study evaluated ChatGPT-4o and Claude 3.5 Sonnet for generating obstetrics and gynecology SCT items on five primary-care diagnostic topics. Sixteen obstetrics and gynecology residents rated 10 AI-generated SCT items against 11 criteria; overall agreement that items met criteria was 90.57% for ChatGPT-4o and 91.48% for Claude 3.5 Sonnet.
Related subquestions: What can AI do for generating script concordance tests?
2024Other
Yavuz Selim Kıyak, Emre Emekli · Revista Española de Educación Médica
This study tested whether ChatGPT-4 and Claude 3 Sonnet could generate Script Concordance Test items for abdominal radiology using a detailed prompt. Sixteen radiologists judged the items; most rated them as assessing clinical reasoning rather than factual recall, and ChatGPT-4 had higher acceptability than Claude across the reported indicators.
Related subquestions: What can AI do for generating script concordance tests?