Is AI the future of evaluation in medical education?? AI vs. human evaluation in objective structured clinical examination
Murat Tekin, Mustafa Onur Yurdal, Çetin Toraman, Güneş Korkmaz, İbrahim Uysal. Is AI the future of evaluation in medical education?? AI vs. human evaluation in objective structured clinical examination. BMC Medical Education. 2025. https://doi.org/10.1186/S12909-025-07241-4
Abstract or concise summary
This cross-sectional study examined whether two multimodal AI systems, ChatGPT-4o and Gemini Flash 1.5, could evaluate medical students’ OSCE performances consistently with human assessors. Conducted at a Turkish state university, it used end-of-year 2023–2024 video recordings from 196 pre-clinical students performing intramuscular injection, square knot tying, basic life support, and urinary catheterization. Each performance was scored with standardized checklists by one real-time human evaluator, two video-based expert evaluators, and the two AI models. Across most skills, AI gave higher and less variable scores than humans. Agreement was strongest for clearly visible procedural steps and weaker for auditory, communication-based, or nuanced fine-motor criteria. Inter-rater reliability was generally low across criteria, and Bland-Altman analyses showed systematic AI scoring bias, especially for knot tying and injection. The authors conclude that AI may support OSCE assessment, but requires refinement, validation, and human oversight before high-stakes use in similar settings or formal grading decisions alone.
DOI: 10.1186/S12909-025-07241-4
URL: https://doi.org/10.1186/s12909-025-07241-4
Journal: BMC Medical Education
Authors: Murat Tekin, Mustafa Onur Yurdal, Çetin Toraman, Güneş Korkmaz, İbrahim Uysal
Last updated: July 5, 2026
Article interpretation
- Main contribution
- This cross-sectional study compared OSCE scores from three human evaluators and two AI systems, ChatGPT-4o and Gemini Flash 1.5, for 196 pre-clinical medical students performing four clinical skills. AI scores were generally higher than human scores, with better alignment on visually observable steps than on auditory or communication-dependent criteria.
- Practical use
- The findings support using AI only as a supplementary OSCE tool, especially for checklist items based on clear visual actions. It may help flag performance patterns or support feedback, but human oversight remains necessary.
- Main caution
- AI tended to over-score students and showed weak agreement with human evaluators across many criteria. It should not be used alone for high-stakes OSCE decisions.
- Research gap
- Further studies should test trained AI models, larger multi-institutional samples, more diverse clinical skills, and stronger multimodal handling of speech, context, and subtle procedural actions.
Collaboration
Working on a study that could address this gap?
If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.
Curator note
This is primary research. The study used end-of-year 2023–2024 OSCE videos from one Turkish state university, with skill-specific samples of 43–58 students. But it is the first in using AI for video scoring in the literature.
Caution
Related links
Related main question areas: What can AI do for automated scoring?
Related subquestions: What can AI do for scoring videos?
Related published prompts: None linked
Tags: