Abstract / Summary
This scoping review maps generative AI applications in workplace-based assessment (WBA) and analyzes validity evidence through Downing’s framework. Following JBI methodology and PRISMA-ScR guidelines, four databases were searched (2022-February 2026) with dual-AI screening and human adjudication. Data were mapped to Downing’s five validity sources using AI-assisted extraction with human verification. Thirteen studies (2024–2025) met inclusion criteria. All AI applications operated on pre-existing text; feedback analysis predominated (9/13). Content (12/13), Response Process (13/13), and Relationship to Other Variables (12/13) were well addressed, while Internal Structure (2/13) and Consequences (4/13) were neglected. AI reasoning transparency, inter-model agreement, and internal consistency were absent (0/13). Current evidence concentrates on AI-human agreement, neglecting reproducibility, bias testing, and learner impact evidence essential for responsible deployment. Realising AI’s potential to reduce faculty burden will require non-supervisor oversight mechanisms for continuous validity monitoring, particularly for direct learner-facing applications.