Abstract / Summary
Abstract Objective Structured Clinical Examinations (OSCEs) are typically scored by a single examiner per station, exposing results to hard-to-audit examiner variability. We developed the Rater-Aware Verification Network (RAVEN), a multimodal Artificial Intelligence (AI) system fusing egocentric video, examiner verbalisations and marks, and virtual reality (VR) action logs as an examiner-conditioned second marker for paediatric OSCEs (retrospective evaluation; 120 students, 442 ratings, eight domains). A confidence-gated hybrid improved agreement with leave-one-examiner-out consensus on the pass/fail decision (AC1: +6.2 percentage points; p = 0.0004) and domain scores (mean AC2: +3.0 points, five of eight significant), with the largest gains on borderline-fail cases (+16.3 points of concordance with consensus). Comparing AI-inferred, rubric-specified, and examiner-articulated criteria revealed implicit practices absent from marking guidelines. Agreement is measured against a panel-derived reference rather than an external ground truth; the system is intended as an audit and flagging aid for human adjudication rather than an autonomous decision-maker.