Abstract / Summary
Background: Serial interpretation of lung volumes requires assessing both reliability and agreement. We examined these properties in a historical clinical dataset. Methods: This retrospective secondary analysis included 72 participants, excluding seven same-session pre/post-bronchodilator pairs. The primary analysis used 19 separately dated pairs labelled pre-bronchodilator at both examinations; the expanded 65-pair analysis was exploratory. Documented intervals were two to three days; the investigator retrospectively confirmed separate-day examinations within six days for all 65 pairs. Total lung capacity (TLC) and residual volume (RV) were principal outcomes; functional residual capacity was secondary and available in the same 19 participants. Analyses included two-way mixed-effects, absolute-agreement, single-measure intraclass correlation coefficients (ICCs), Lin’s concordance correlation coefficient, Gaussian reference limits, empirical difference quantiles, and leave-one-out analyses. Results: Fourteen primary participants had recorded chronic obstructive pulmonary disease labels without independently verified diagnostic post-bronchodilator spirometry. Primary TLC and RV ICCs were 0.947 (95% bootstrap confidence interval 0.838–0.988) and 0.940 (0.818–0.974), respectively. Mean second-minus-first differences were −0.033 and −0.015 L, with Gaussian reference limits of −0.964 to 0.898 L and −0.934 to 0.904 L, respectively. Corresponding empirical 2.5th–97.5th percentiles were −1.122 to 0.641 L and −0.875 to 0.803 L. Removing one observation reduced the standard deviation of paired differences by up to 28.8% for TLC and 14.8% for RV. Conclusions: High reliability estimates coexisted with appreciable paired variation. Unverified technical quality and clinical stability preclude attributing this variation solely to measurement error or establishing clinical response thresholds.