Automated scoring of student videos in medical education: a comparison between a large language model and expert evaluation
Yavuz Selim Kıyak, Özlem Ülkü Bulut, Özlem Coşkun, Işıl İrem Budakoğlu · Journal of Microbiology & Biology Education
This study shows that Gemini 2.5 Pro did not reliably match expert scoring of 139 medical student EBM video presentations. A rubric-only prompt over-scored students, while a “critical expert” prompt under-scored them, showing that prompt strategy changed both the size and direction of scoring bias.
Related subquestions: What can AI do for scoring videos?