AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory
2024Original researchOther

Using Large Language Models to Generate Script Concordance Test in Medical Education: ChatGPT and Claude

Yavuz Selim Kıyak, Emre Emekli. Using Large Language Models to Generate Script Concordance Test in Medical Education: ChatGPT and Claude. Revista Española de Educación Médica. 2024. https://doi.org/10.6018/EDUMED.636331

Abstract or concise summary

We aimed to determine the quality of AI-generated (ChatGPT-4 and Claude 3) Script Concordance Test (SCT) items through an expert panel. We generated SCT items on abdominal radiology using a complex prompt in large language model (LLM) chatbots (ChatGPT-4 and Claude 3 (Sonnet) in April 2024) and evaluated the items’ quality through an expert panel of 16 radiologists. Expert panel, which was blind to the origin of the items provided without modifications, independently answered each item and assessed them using 12 quality indicators. Data analysis included descriptive statistics, bar charts to compare responses against accepted forms, and a heatmap to show performance in terms of the quality indicators. SCT items generated by chatbots assess clinical reasoning rather than only factual recall (ChatGPT: 92.50%, Claude: 85.00%). The heatmap indicated that the items were generally acceptable, with most responses favorable across quality indicators (ChatGPT: 71.77%, Claude: 64.23%). The comparison of the bar charts with acceptable and unacceptable forms revealed that 73.33% and 53.33% of the questions in the items can be considered acceptable, respectively, for ChatGPT and Claude. The use of LLMs to generate SCT items can be helpful for medical educators by reducing the required time and effort. Although the prompt provides a good starting point, it remains crucial to review and revise AI-generated SCT items before educational use. The prompt and the custom GPT, “Script Concordance Test Generator”, available at https://chatgpt.com/g/g-RlzW5xdc1-script-concordance-test-generator, can streamline SCT item development.

DOI: 10.6018/EDUMED.636331

URL: https://doi.org/10.6018/edumed.636331

Journal: Revista Española de Educación Médica

Authors: Yavuz Selim Kıyak, Emre Emekli

Last updated: July 3, 2026

Article interpretation

Main contribution
This study tested whether ChatGPT-4 and Claude 3 Sonnet could generate Script Concordance Test items for abdominal radiology using a detailed prompt. Sixteen radiologists judged the items; most rated them as assessing clinical reasoning rather than factual recall, and ChatGPT-4 had higher acceptability than Claude across the reported indicators.
Practical use
Medical educators can use LLMs to create first drafts of SCT items. Use the prompt template.
Main caution
AI-generated SCT items still need expert review and revision before educational use. Some items had unacceptable response patterns, including patterns resembling single-best-answer questions or non-discriminatory items.
Research gap
The evidence comes from 10 SCT items, five abdominal radiology pain presentations, two proprietary LLMs, and one single-center expert panel. Further studies should test more specialties, more learner groups, newer and open-source models, human-written comparators, and learner-response validity evidence.

Collaboration

Working on a study that could address this gap?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Curator note

This is the first primary evaluation study. Its useful contribution is the combination of a detailed SCT prompt, blinded expert-panel review, 12 quality indicators, and response-pattern analysis against acceptable and unacceptable SCT forms.

Caution

Do not consider this study as an ultimate evidence that LLM-generated SCTs are valid for summative assessment. It shows potential for generating drafts, but item quality varied and the study did not test performance with students or compare against human-written SCTs.

Prompt / method

Includes prompt
Yes

Related links