AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory
2026Original researchOther

Large language models in professionalism and ethical decision-making education: a transferable process for generating concordance of professional judgment items

İş-Kara T, Kıyak YS. Large language models in professionalism and ethical decision-making education: a transferable process for generating concordance of professional judgment items. Academic Medicine. Published online May 9, 2026. doi:10.1093/acamed/wvag147.

Abstract or concise summary

Concordance-based tools, such as the Concordance of Professional Judgment (CoPJ) tool, allow students to practice ethical decision-making and professionalism under uncertain conditions, but developing high-quality CoPJ items (including writing vignettes and convening expert panels) is resource intensive. Thus, programs often lack the practical infrastructure needed to integrate the CoPJ into curricula at the scale necessary for repeated learner practice. Large language models (LLMs) may help address this burden, but a transferable process to generate CoPJ items aligned with local ethics frameworks and expert oversight is needed. This study implemented and evaluated a transferable process for generating CoPJ items using general purpose LLMs. Using a generic zero-shot prompt, ChatGPT-5 Thinking and Gemini 2.5 Pro were each prompted to generate 1 CoPJ item for all 28 topics in the Turkish Medical Association’s ethical declarations (56 items total) in August 2025. Between August and September 2025, 2 faculty reviewers independently rated each LLM-generated item using a 20-criterion checklist and a 4-level global suitability judgment (accepted, minor revision, major revision, rejected), with adjudication as needed. Gemini 2.5 Pro achieved 96.4% overall global suitability, with 27/28 items judged to be suitable (accepted or minor revision). ChatGPT-5 Thinking’s overall global suitability was 89.3% (25/28 items judged to be suitable). Gemini 2.5 Pro had an overall mean checklist success score of 97.5% (19.5/20), and ChatGPT-5 Thinking had a score of 89.0% (17.8/20). Future work will examine how LLM-generated CoPJ items function in real learning environments, including a planned learner-based comparison with human-written items. Additional studies across languages, professional programs, and cultural contexts will further help determine the broader adaptability of this process.

DOI: 10.1093/ACAMED/WVAG147

URL: https://doi.org/10.1093/acamed/wvag147

Journal: Academic Medicine

Authors: Tuğba İş-Kara, Yavuz Selim Kıyak

Last updated: July 19, 2026

Article interpretation

Main contribution
This innovation report presents an evaluated human-in-the-loop workflow for generating Concordance of Professional Judgment items from a locally defined ethics framework. ChatGPT-5 Thinking and Gemini 2.5 Pro each generated 28 Turkish items; after independent faculty review and adjudication, 25/28 and 27/28 items, respectively, were accepted or required only minor revision.
Practical use
Educators can define local ethics topics and learner parameters, use the structured zero-shot prompt to generate first drafts, and then apply the study's explicit checklist, independent faculty review, and adjudication process. The demonstrated use is formative content development; generated items should be revised and locally approved before learner use.
Main caution
The evaluation used faculty judgments rather than learner outcomes and did not compare the items with human-written CoPJ items. The study was also limited to two models, one prompt template, and one linguistic and cultural context.
Research gap
Future studies need learner-based testing, comparison with human-written CoPJ items, response-process and validity evidence, and evaluation across languages, professional programs, cultural contexts, prompts, and models.

Collaboration

Working on a study that could address this gap?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Curator note

The strongest contribution is the complete content-development process: local framework alignment, a reusable prompt, comprehensive topic coverage, a 20-criterion quality checklist, independent review, and adjudication. Interpret the results as evidence of drafting feasibility and faculty-rated suitability, not educational effectiveness.

Caution

Do not use LLM-generated CoPJ items without expert review. Check factual accuracy, local ethical and legal relevance, ambiguity, learner level, and the depth of explanations before implementation.

Prompt / method

Includes prompt
Yes
Original prompt status
Full prompt published in the article's official Supplementary Material 1 under CC BY 4.0.
Prompt purpose
Generate formative CoPJ item drafts for health-profession learners that represent authentic ethical or professional uncertainty and make divergent reasoning visible.
Prompt summary
A reusable zero-shot template parameterized by topic, learner level, number of items, number of experts, locale or policy hooks, and language. It requests a brief vignette, one behavior or attitude, a forced four-point scale, a panel distribution, short explanations, an educational synthesis, action reminders, tags, and a quality checklist.
Required inputs
Topic, Learner level, Number of items, Number of experts, Locale or policy hooks, Language

Related links