Abstract / Summary
BackgroundA primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification--requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners.
ObjectiveTo quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it.
MethodsIn a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss {kappa}, Krippendorffs , Gwets AC1, and ICC(2,1)/ICC(2,k).
ResultsHuman item agreement was moderate ({kappa} 0.409-0.467), as was single-rater reliability (ICC(2,1) 0.539-0.694), against good panel reliability (ICC(2,k) 0.854-0.919). Recorded pass/fail agreement was fair ({kappa} 0.297-0.388) and critical-failure agreement near zero for two skills ({kappa} 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills.
ConclusionsHuman evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration--a baseline, not a limit.