Abstract / Summary
Physical activity (PA) during childhood is essential for healthy growth and development. Device-based measures provide accurate PA estimates but are costly and resource-intensive, leading to reliance on self-report instruments where device-based assessments are not feasible. However, these instruments have been less extensively validated among children and youth in South Asia relative to their Western and higher-income counterparts. The primary objectives of this study were to evaluate the criterion and concurrent validity of the modified Youth Physical Activity Questionnaire (YPAQ) against accelerometer-derived estimates considered as the criterion measure for assessing moderate-to-vigorous physical activity (MVPA) and to assess its test-retest reliability. The secondary objective was to evaluate the moderating effects of sex, age, day of week, and weight status on validity and reliability coefficients with particular attention to potential gender-related differences in measurement performance. This cross-sectional validation study was conducted among school-going healthy children aged 9–14 years recruited from low-resource settings of Karachi. Participants with any physical or mental disabilities were excluded. Device-based physical activity was assessed using GT3X accelerometers worn for seven consecutive days. The proxy-assisted YPAQ was administered twice, one week apart. Criterion validity was evaluated using correlation coefficients (r) and test-retest reliability with ICCs and 95% confidence intervals. Agreement was quantified using Bland-Altman analysis, estimating mean bias and 95% limits of agreement as percent differences. Proportional bias was assessed by regressing percent difference on the mean of the two methods; a slope with 95% confidence intervals excluding zero was interpreted as evidence of systematic variation across the MVPA range. Of the 252 enrolled children, 234 (93%) provided valid accelerometer data. The criterion validity coefficient for MVPA was modest [ r = 0.37, (95% CI: 0.29 to 0.44), r² = 0.137] and test-retest reliability fell in the moderate range [ICC = 0.73, (95% CI: 0.68 to 0.77)]. Bland-Altman plots detected a mean bias of [+15.06%, (95% CI: -150.05% to 180.17%)] indicating weak concordance and systematic overestimation by YPAQ with a significant proportional bias [slope = 0.053, (95% CI: 0.046 to 0.061), p < 0.001]. Sex-stratified validity was even weaker: boys [ r = 0.25, (95% CI: 0.13 to 0.36), r²=0.063] and girls [ r = 0.28, (95% CI: 0.15 to 0.40), r² =0.078]. Stratified Bland Altman demonstrated greater bias in boys [26.22%, (95% CI: -141.79 to 194.23%)] compared to girls [15.50%, (95% CI: -158.30 to 189.29)] with a significant proportional comparable bias between both sexes, [boys’ slope = 0.055, (95% CI: 0.046 to 0.064), p < 0.001] and [girls’ slope = 0.053, (95% CI: 0.039 to 0.065), p < 0.001]. Overall reliability was moderate and consistent across boys and girls. The YPAQ demonstrated acceptable test-retest reliability but limited criterion validity against accelerometer-derived measures with modest agreement and systematic overestimation of MVPA. Proportional bias was evident with underestimation at lower activity levels and progressive overestimation at higher levels, particularly among boys. These findings indicate that YPAQ is better suited to group-level monitoring than individual-level assessment. However, proportional and sex-specific bias limit its validity in its current form, underscoring the need for population-specific calibration before broader application in research or surveillance settings.