Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude
Zuhal Yapıcı Coşkun, Yavuz Selim Kıyak, Özlem Coşkun, Işıl İrem Budakoğlu, Özhan Özdemir. Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude. Medical Teacher. 2025. https://doi.org/10.1080/0142159X.2025.2497888
DOI: 10.1080/0142159X.2025.2497888
URL: https://doi.org/10.1080/0142159x.2025.2497888
Journal: Medical Teacher
Authors: Zuhal Yapıcı Coşkun, Yavuz Selim Kıyak, Özlem Coşkun, Işıl İrem Budakoğlu, Özhan Özdemir
Last updated: July 3, 2026
Article interpretation
- Main contribution
- This cross-sectional study evaluated ChatGPT-4o and Claude 3.5 Sonnet for generating obstetrics and gynecology SCT items on five primary-care diagnostic topics. Sixteen obstetrics and gynecology residents rated 10 AI-generated SCT items against 11 criteria; overall agreement that items met criteria was 90.57% for ChatGPT-4o and 91.48% for Claude 3.5 Sonnet.
- Practical use
- Educators can use these LLMs to draft SCT items for clinical reasoning assessment. Use the prompt template.
- Main caution
- The weakest criterion for both models was appropriate difficulty for medical students: 71.25% for ChatGPT-4o and 76.25% for Claude 3.5 Sonnet. This means the models may generate plausible SCTs that still need adjustment to match learner level.
- Research gap
- The study did not compare AI-generated SCTs with human-written SCTs, did not test the items with students, and covered only five diagnostic topics in obstetrics and gynecology. Further work should examine other specialties, non-diagnostic SCTs, reasoning-focused models, human comparators, and validity evidence from learner responses.
Collaboration
Working on a study that could address this gap?
If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.
Curator note
This is the second evaluation study of AI-generated SCT item quality, not evidence that the items improve learning or assessment outcomes. Its main value is showing that newer models, ChatGPT-4o and Claude 3.5 Sonnet.
Caution
Related links
Related main question areas: What can AI do for generating assessment materials?
Related subquestions: What can AI do for generating script concordance tests?
Related published prompts: Script Concordance Test generation prompt template in medical education
Tags: script concordance test