AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory
2025

Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude

Zuhal Yapıcı Coşkun, Yavuz Selim Kıyak, Özlem Coşkun, Işıl İrem Budakoğlu, Özhan Özdemir. Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude. Medical Teacher. 2025. https://doi.org/10.1080/0142159X.2025.2497888

DOI: 10.1080/0142159X.2025.2497888

URL: https://doi.org/10.1080/0142159x.2025.2497888

Journal: Medical Teacher

Authors: Zuhal Yapıcı Coşkun, Yavuz Selim Kıyak, Özlem Coşkun, Işıl İrem Budakoğlu, Özhan Özdemir

Last updated: July 3, 2026

Article interpretation

Main contribution
This cross-sectional study evaluated ChatGPT-4o and Claude 3.5 Sonnet for generating obstetrics and gynecology SCT items on five primary-care diagnostic topics. Sixteen obstetrics and gynecology residents rated 10 AI-generated SCT items against 11 criteria; overall agreement that items met criteria was 90.57% for ChatGPT-4o and 91.48% for Claude 3.5 Sonnet.
Practical use
Educators can use these LLMs to draft SCT items for clinical reasoning assessment. Use the prompt template.
Main caution
The weakest criterion for both models was appropriate difficulty for medical students: 71.25% for ChatGPT-4o and 76.25% for Claude 3.5 Sonnet. This means the models may generate plausible SCTs that still need adjustment to match learner level.
Research gap
The study did not compare AI-generated SCTs with human-written SCTs, did not test the items with students, and covered only five diagnostic topics in obstetrics and gynecology. Further work should examine other specialties, non-diagnostic SCTs, reasoning-focused models, human comparators, and validity evidence from learner responses.

Collaboration

Working on a study that could address this gap?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Curator note

This is the second evaluation study of AI-generated SCT item quality, not evidence that the items improve learning or assessment outcomes. Its main value is showing that newer models, ChatGPT-4o and Claude 3.5 Sonnet.

Caution

Do not use this study to justify unsupervised AI item generation or immediate use in summative assessment.

Related links