Abstract / Summary
Background/Objectives: Catheter-appropriateness audits report a single rate without a measure of its reliability. We measured how far an emergency department (ED) estimate depends on the clinician applying the checklist, and which items carry that dependence. Methods: We enrolled 117 consecutive adults catheterized in a tertiary ED between 21 January and 21 May 2026. Three emergency physicians independently classified every patient against a ten-item checklist drawn from the 2009 CDC/HICPAC guideline and the Ann Arbor criteria. A designated senior rater served as comparator, a clinical judgment rather than a criterion standard. We computed standalone rates, per-item Fleiss’ κ and Gwet’s AC1, patient-level agreement, and a post hoc counterfactual item-replacement analysis. Results: Ten patients (8.5%, 95% CI 4.7–15.0) had no valid indication according to the senior rater. Auditing alone, the two index raters would have reported 0.9% and 25.6%. Pooled agreement over 1170 patient-by-indication decisions was substantial by κ (0.630 and 0.707) and almost perfect by AC1 (0.885 and 0.925); excluding the complementary item changed neither. Patient-level agreement was lower (three-rater κ 0.199, AC1 0.792). Disagreement concentrated in hourly urine-output monitoring (AC1 0.617) and neurogenic bladder, endorsed in 44 patients by one rater and none by another (AC1 0.668). That item accounted for 48% of one rater’s excess endorsements; replacing that rater’s item-5 ratings with the senior rater’s narrowed the between-rater range from 30-fold to 6-fold. Conclusions: In this cohort the audit rate was rater-dependent, and the dependence concentrated in identifiable checklist items. Item-level operational definitions and prospective calibration should be evaluated as strategies for improving audit reliability, and per-item agreement should be reported with a prevalence-robust coefficient alongside κ.