AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory
2025Original research

Is AI the future of evaluation in medical education?? AI vs. human evaluation in objective structured clinical examination

Murat Tekin, Mustafa Onur Yurdal, Çetin Toraman, Güneş Korkmaz, İbrahim Uysal. Is AI the future of evaluation in medical education?? AI vs. human evaluation in objective structured clinical examination. BMC Medical Education. 2025. https://doi.org/10.1186/S12909-025-07241-4

Abstract or concise summary

This cross-sectional study examined whether two multimodal AI systems, ChatGPT-4o and Gemini Flash 1.5, could evaluate medical students’ OSCE performances consistently with human assessors. Conducted at a Turkish state university, it used end-of-year 2023–2024 video recordings from 196 pre-clinical students performing intramuscular injection, square knot tying, basic life support, and urinary catheterization. Each performance was scored with standardized checklists by one real-time human evaluator, two video-based expert evaluators, and the two AI models. Across most skills, AI gave higher and less variable scores than humans. Agreement was strongest for clearly visible procedural steps and weaker for auditory, communication-based, or nuanced fine-motor criteria. Inter-rater reliability was generally low across criteria, and Bland-Altman analyses showed systematic AI scoring bias, especially for knot tying and injection. The authors conclude that AI may support OSCE assessment, but requires refinement, validation, and human oversight before high-stakes use in similar settings or formal grading decisions alone.

DOI: 10.1186/S12909-025-07241-4

URL: https://doi.org/10.1186/s12909-025-07241-4

Journal: BMC Medical Education

Authors: Murat Tekin, Mustafa Onur Yurdal, Çetin Toraman, Güneş Korkmaz, İbrahim Uysal

Last updated: July 5, 2026

Article interpretation

Main contribution
This cross-sectional study compared OSCE scores from three human evaluators and two AI systems, ChatGPT-4o and Gemini Flash 1.5, for 196 pre-clinical medical students performing four clinical skills. AI scores were generally higher than human scores, with better alignment on visually observable steps than on auditory or communication-dependent criteria.
Practical use
The findings support using AI only as a supplementary OSCE tool, especially for checklist items based on clear visual actions. It may help flag performance patterns or support feedback, but human oversight remains necessary.
Main caution
AI tended to over-score students and showed weak agreement with human evaluators across many criteria. It should not be used alone for high-stakes OSCE decisions.
Research gap
Further studies should test trained AI models, larger multi-institutional samples, more diverse clinical skills, and stronger multimodal handling of speech, context, and subtle procedural actions.

Collaboration

Working on a study that could address this gap?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Curator note

This is primary research. The study used end-of-year 2023–2024 OSCE videos from one Turkish state university, with skill-specific samples of 43–58 students. But it is the first in using AI for video scoring in the literature.

Caution

The results may not generalize beyond the study setting, skills, prompts, video setup, or model versions. Newer AI models could provide better results.

Related links

Related main question areas: What can AI do for automated scoring?

Related subquestions: What can AI do for scoring videos?

Related published prompts: None linked

Tags: