AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory

What can AI do for generating script concordance tests?

AI can produce rapid first drafts of Script Concordance Test (SCT) items when given a structured prompt, target learner level, clinical topic, and assessment focus. AI should not be treated as an autonomous SCT writer. The current evidence supports human-in-the-loop SCT development, not unsupervised use in summative or high-stakes assessment.

Evidence summary

Two studies evaluated LLM-generated SCTs with expert panels. In abdominal radiology, ChatGPT-4 and Claude 3 generated 10 SCT items; 16 radiologists judged most as assessing clinical reasoning, but only 73.33% of ChatGPT questions and 53.33% of Claude questions had acceptable response patterns. In obstetrics and gynecology, ChatGPT-4o and Claude 3.5 Sonnet generated 10 SCT items; 16 residents rated overall quality highly, with 90.57% and 91.48% agreement that items met criteria. The main weakness was matching difficulty to medical students.

Practical starting points

  • Start with a blueprint, learner level, clinical setting, topic, and decision focus.
  • Ask AI to generate an item by using the prompt template.
  • Have experts review and revise the items.

Last updated: July 3, 2026

Collaboration

Planning a study related to this question?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this question or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Related articles

Selected studies

2025

Large language models for generating script concordance test in obstetrics and gynecology: ChatGPT and Claude

Zuhal Yapıcı Coşkun, Yavuz Selim Kıyak, Özlem Coşkun, Işıl İrem Budakoğlu, Özhan Özdemir · Medical Teacher

This cross-sectional study evaluated ChatGPT-4o and Claude 3.5 Sonnet for generating obstetrics and gynecology SCT items on five primary-care diagnostic topics. Sixteen obstetrics and gynecology residents rated 10 AI-generated SCT items against 11 criteria; overall agreement that items met criteria was 90.57% for ChatGPT-4o and 91.48% for Claude 3.5 Sonnet.

Related subquestions: What can AI do for generating script concordance tests?

2024Other

Using Large Language Models to Generate Script Concordance Test in Medical Education: ChatGPT and Claude

Yavuz Selim Kıyak, Emre Emekli · Revista Española de Educación Médica

This study tested whether ChatGPT-4 and Claude 3 Sonnet could generate Script Concordance Test items for abdominal radiology using a detailed prompt. Sixteen radiologists judged the items; most rated them as assessing clinical reasoning rather than factual recall, and ChatGPT-4 had higher acceptability than Claude across the reported indicators.

Related subquestions: What can AI do for generating script concordance tests?

Published prompts

Related prompt architectures

Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License

Script Concordance Test generation prompt template in medical education

Prompt

Related subquestions: What can AI do for generating script concordance tests?

Open prompt page

Practical guides

Linked guides