AI MedEd ObservatoryWhat Can AI Do in Medical Education? An observatory

Evidence, prompts, checklists, and practical interpretation for assessment, feedback, scoring, and teaching materials.

Mantra: Evidence before automation. Judgment before adoption.Base path /ai-meded-observatory
2026Review articleScoping review

Applications and Outcomes of Large‑Language‑Model‑Generated Feedback in Undergraduate Medical Education: A Scoping Review

Yavuz Selim Kıyak, Tuğba İş-Kara, Emre Emekli. Applications and Outcomes of Large‑Language‑Model‑Generated Feedback in Undergraduate Medical Education: A Scoping Review. Medical Science Educator. 2026. https://doi.org/10.1007/S40670-025-02621-3

Abstract or concise summary

This scoping review maps LLM-generated feedback in undergraduate medical education. Authors searched PubMed and Web of Science through 10 September 2025, screened 4325 records, and included 42 studies. Most were from Global North countries and used OpenAI GPT models. Feedback was mainly used in simulated clinical encounters and text-based assessment tasks, including MCQs, clinical reasoning exercises, history-taking, communication practice, and writing. Evidence was limited: 22 studies had no student data, 10 measured student reactions, 10 measured learning, and only 8 used randomized trials. Reported benefits were mostly perceived usefulness or short-term gains, sometimes comparable to expert feedback. Accuracy varied, and one study found AI errors could shape student decisions despite peer discussion. No study assessed clinical behavior or patient-level outcomes. The authors recommend supervised, low-stakes use with safeguards, not unsupervised high-stakes implementation. Review limits include English-only peer-reviewed studies and no risk-of-bias appraisal.

DOI: 10.1007/S40670-025-02621-3

URL: https://doi.org/10.1007/s40670-025-02621-3

Journal: Medical Science Educator

Authors: Yavuz Selim Kıyak, Tuğba İş-Kara, Emre Emekli

Last updated: July 5, 2026

Article interpretation

Main contribution
This scoping review maps 42 peer-reviewed studies on LLM-generated feedback in undergraduate medical education, showing that LLMs are being used across MCQs, clinical reasoning, simulated patients, communication skills, academic writing, and self-directed learning. Evidence mainly supports feasibility and short-term learning or learner reaction, not behavior change or patient outcomes.
Practical use
Educators can use this review to identify where LLM feedback has been trialled, what forms it has taken, and which outcomes have been measured. It supports cautious use of LLM feedback as a supervised adjunct for timely, personalized formative feedback.
Main caution
The evidence remains preliminary. Many studies had no student outcome data, relied on informal or expert-only evaluation, and most used GPT-family models in high-income settings. Findings should not be generalized to all models, settings, or learners.
Research gap
Future studies need to test whether LLM feedback changes learner behavior in clinical settings, affects patient care, works across lower-resource contexts, and remains accurate across different models. The review also calls for clearer reporting of prompts and model settings, bias and accessibility audits, cost-effectiveness work, and implementation studies.

Collaboration

Working on a study that could address this gap?

If you have a mature study idea, drafted protocol, or ethics-stage project related to this research gap or another gap in AI and medical education, you can share a collaboration proposal. Please include enough detail to assess fit, feasibility, and possible collaboration. Not every proposal will lead to collaboration.

Share a collaboration proposal

Curator note

This is a scoping review, it synthesizes where and how LLM feedback has been used rather than proving effectiveness. The search was limited to English-language peer-reviewed literature and cut off on 10 September 2025. It is the most up-to-date review as of July 2026.

Caution

Do not treat LLM feedback as equivalent to faculty feedback across contexts. The article recommends careful supervision and clear governance to safeguard accuracy, equity, and learner trust.

Prompt / method

Includes prompt
No

Related links

Related main question areas: What can AI do for feedback?

Related subquestions: What can AI do for generating feedback?

Related published prompts: None linked

Tags: feedback, scoping review, undergraduate