AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory
2026Original researchOther

Automated scoring of student videos in medical education: a comparison between a large language model and expert evaluation

Yavuz Selim Kıyak, Özlem Ülkü Bulut, Özlem Coşkun, Işıl İrem Budakoğlu. Automated scoring of student videos in medical education: a comparison between a large language model and expert evaluation. Journal of Microbiology & Biology Education. 2026. https://doi.org/10.1128/JMBE.00010-26

Abstract or concise summary

Video-based assignments are used in medical education, yet expert scoring is time-intensive. Large language models (LLMs) offer scalable alternatives, but their validity for evaluating multimodal student work is uncertain. We examined whether a state-of-the-art LLM (Gemini 2.5 Pro) could approximate expert scoring of medical student video presentations in evidence-based medicine. A total of 139 student submissions were evaluated by an experienced faculty member (reference standard) and by the LLM under two prompting strategies: (i) rubric-only and (ii) critical expert-style. Scores across 12 rubric items (maximum 95 points) were compared using paired t -tests, effect sizes, and Bland-Altman analysis. Expert scoring yielded a mean total of 62.0 (SD 13.6). The rubric-only prompt systematically overestimated performance (mean 87.7, SD 8.0, P < 0.001; bias –25.7). The critical prompt produced lower scores (mean 53.5, SD 10.4, P < 0.001; bias +8.5). At the item level, rubric-only prompting aligned better with mechanical tasks (e.g., keywords and referencing), whereas the critical prompt penalized appraisal and synthesis disproportionately. Prompting strategy substantially influenced LLM scoring, generating opposite biases relative to expert evaluation. The novel contribution of this study is that prompt strategy can alter not only the magnitude but also the direction of scoring bias. Calibration approaches, such as context engineering, may help align AI scoring with expert judgment. While AI-generated feedback shows promise for formative assessment, reliable summative use requires careful validation.

DOI: 10.1128/JMBE.00010-26

URL: https://doi.org/10.1128/jmbe.00010-26

Journal: Journal of Microbiology & Biology Education

Authors: Yavuz Selim Kıyak, Özlem Ülkü Bulut, Özlem Coşkun, Işıl İrem Budakoğlu

Last updated: June 28, 2026

Article interpretation

Main contribution
This study shows that Gemini 2.5 Pro did not reliably match expert scoring of 139 medical student EBM video presentations. A rubric-only prompt over-scored students, while a “critical expert” prompt under-scored them, showing that prompt strategy changed both the size and direction of scoring bias.
Practical use
Educators can use this article to justify cautious, locally validated use of LLMs for formative feedback on video assignments, especially for more mechanical rubric items such as keywords, referencing, structure, and clarity.
Main caution
The findings do not support using an LLM as an interchangeable summative rater. Scores varied substantially by prompt, and the model was least aligned with expert judgment on appraisal, synthesis, and other judgment-heavy tasks.
Research gap
Future work needs multiple expert raters, larger and more diverse settings, more assessment formats, comparisons across models and prompts, formal analysis of feedback quality, and tested calibration methods before routine grading use.

Collaboration

Working on a study that could address this gap?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Curator note

Useful evidence for medical educators considering AI-assisted assessment of student videos. Its main value is not that AI can grade videos, but that seemingly reasonable prompts can produce opposite scoring errors.

Caution

Interpret the results as preliminary: the study used one institution, one assessment task, one LLM, two prompt styles, and a single expert reference standard; 152 of 291 submissions were excluded from analysis for consent, slide, or technical reasons.

Prompt / method

Includes prompt
Yes

Related links

Related main question areas: What can AI do for automated scoring?, What can AI do for feedback?

Related subquestions: What can AI do for scoring videos?

Related published prompts: None linked

Tags: video scoring, validity, calibration