Abstract / Summary
Rapid identification of adults with poor physical fitness may help target early behavioural and clinical resources in resource-constrained high-altitude settings. We developed and internally–externally validated (leave-one-city-out) interpretable machine-learning classifiers to identify adults with currently poor physical fitness using Tibet National Physical Fitness Monitoring data. This cross-sectional analysis included 5,750 adults aged 20–69 years from Nyingchi (3100 m), Lhasa (3658 m), and Nagqu (4500 m), drawn from 6,160 sampled residents. Poor physical fitness was defined using age-specific total-score cut-points from the Chinese National Physical Fitness Standard (< 23 points at age 20–39 years, < 18 at 40–59 years, and < 15 at 60–69 years). To limit outcome (label) leakage, predictors were prespecified to variables obtainable before the scored performance battery: demographics, exercise behaviour, adiposity markers, and resting haemodynamic measurements. Logistic regression, random forest, and XGBoost were compared using internal-external cross-validation by site (leave-one-city-out). Pooled held-out predictions were used to estimate discrimination, precision-recall, calibration, classification metrics at clinically relevant thresholds, and decision-curve performance, with 95% confidence intervals from bootstrap resampling. Robustness was examined with multiple imputation, an anthropometry-free model, and site-adjusted analyses. SHAP applied to a matched gradient-boosting model was used to describe the dominant predictors, and concordance with random-forest permutation importance was assessed. Poor physical fitness was present in 2,046 of 5,750 adults (35.6%). Event rates varied markedly by site, from 27.6% in Lhasa to 46.9% in Nyingchi. The random forest showed the highest pooled held-out discrimination point estimates (AUROC 0.664, 95% CI 0.649–0.681; AUPRC 0.560, 95% CI 0.537–0.583; Brier score 0.212, 95% CI 0.208–0.216), followed by XGBoost (AUROC 0.652, 95% CI 0.637–0.667) and logistic regression (AUROC 0.646, 95% CI 0.631–0.661). Differences in discrimination between models were small, and XGBoost had a slightly lower Brier score than the random forest. Calibration of the champion random forest was moderate (intercept − 0.26; slope 0.86). At a prevalence-matched threshold the random forest reached a sensitivity of 0.71 and specificity of 0.48; at a lower fixed threshold of 0.30 sensitivity rose to 0.81 with specificity 0.33, underscoring the screening-only nature of the tool. Results were robust in complete-case analysis, under multiple imputation (AUROC 0.667), and in a leaner predictor set; an anthropometry-free model performed substantially worse (AUROC 0.55). SHAP highlighted BMI, education level, height, sex, arm skinfold thickness, urbanicity, exercise intensity, waist-to-height ratio, and systolic blood pressure as the dominant contributors, and the ordering of the 12 leading predictors showed substantial rank agreement with random-forest permutation importance (Spearman ρ = 0.78). Interpretable tree-based models using simple demographic, behavioural, and anthropometric variables achieved only moderate discrimination for identifying current poor physical fitness in high-altitude adults, with site-wise validation indicating limited but non-trivial transportability. Adiposity and socioeconomic-behavioural features dominated model attribution, although the anthropometric scoring component of the fitness standard partly explains the prominence of body-size variables. As a preliminary screening aid, the model may help prioritise adults for more complete fitness assessment where full testing is difficult, but it should complement rather than replace direct physical fitness assessment and requires independent external validation before use. Not applicable.