Automated scoring of student videos in medical education: a comparison between a large language model and expert evaluation
Yavuz Selim Kıyak, Özlem Ülkü Bulut, Özlem Coşkun, Işıl İrem Budakoğlu. Automated scoring of student videos in medical education: a comparison between a large language model and expert evaluation. Journal of Microbiology & Biology Education. 2026. https://doi.org/10.1128/JMBE.00010-26
Abstract or concise summary
Video-based assignments are used in medical education, yet expert scoring is time-intensive. Large language models (LLMs) offer scalable alternatives, but their validity for evaluating multimodal student work is uncertain. We examined whether a state-of-the-art LLM (Gemini 2.5 Pro) could approximate expert scoring of medical student video presentations in evidence-based medicine. A total of 139 student submissions were evaluated by an experienced faculty member (reference standard) and by the LLM under two prompting strategies: (i) rubric-only and (ii) critical expert-style. Scores across 12 rubric items (maximum 95 points) were compared using paired t -tests, effect sizes, and Bland-Altman analysis. Expert scoring yielded a mean total of 62.0 (SD 13.6). The rubric-only prompt systematically overestimated performance (mean 87.7, SD 8.0, P < 0.001; bias –25.7). The critical prompt produced lower scores (mean 53.5, SD 10.4, P < 0.001; bias +8.5). At the item level, rubric-only prompting aligned better with mechanical tasks (e.g., keywords and referencing), whereas the critical prompt penalized appraisal and synthesis disproportionately. Prompting strategy substantially influenced LLM scoring, generating opposite biases relative to expert evaluation. The novel contribution of this study is that prompt strategy can alter not only the magnitude but also the direction of scoring bias. Calibration approaches, such as context engineering, may help align AI scoring with expert judgment. While AI-generated feedback shows promise for formative assessment, reliable summative use requires careful validation.
DOI: 10.1128/JMBE.00010-26
URL: https://doi.org/10.1128/jmbe.00010-26
Journal: Journal of Microbiology & Biology Education
Authors: Yavuz Selim Kıyak, Özlem Ülkü Bulut, Özlem Coşkun, Işıl İrem Budakoğlu
Last updated: June 28, 2026
Article interpretation
- Main contribution
- This study shows that Gemini 2.5 Pro did not reliably match expert scoring of 139 medical student EBM video presentations. A rubric-only prompt over-scored students, while a “critical expert” prompt under-scored them, showing that prompt strategy changed both the size and direction of scoring bias.
- Practical use
- Educators can use this article to justify cautious, locally validated use of LLMs for formative feedback on video assignments, especially for more mechanical rubric items such as keywords, referencing, structure, and clarity.
- Main caution
- The findings do not support using an LLM as an interchangeable summative rater. Scores varied substantially by prompt, and the model was least aligned with expert judgment on appraisal, synthesis, and other judgment-heavy tasks.
- Research gap
- Future work needs multiple expert raters, larger and more diverse settings, more assessment formats, comparisons across models and prompts, formal analysis of feedback quality, and tested calibration methods before routine grading use.
Collaboration
Working on a study that could address this gap?
If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.
Curator note
Useful evidence for medical educators considering AI-assisted assessment of student videos. Its main value is not that AI can grade videos, but that seemingly reasonable prompts can produce opposite scoring errors.
Caution
Prompt / method
- Includes prompt
- Yes
Related links
Related main question areas: What can AI do for automated scoring?, What can AI do for feedback?
Related subquestions: What can AI do for scoring videos?
Related published prompts: None linked
Tags: video scoring, validity, calibration