Abstract / Summary
Adolescent overweight and obesity are shaped by demographic, behavioral and psychosocial factors, but it remains unclear whether broader predictor information improves cross-sectional classification of current overweight/obesity status more consistently than changing the machine-learning algorithm. This cross-sectional secondary analysis used 2023 Youth Risk Behavior Surveillance System data from 17,814 United States adolescents aged 12–18 years. Survey-aware multivariable logistic regression was used to identify independent correlates, while Random Forest and XGBoost were evaluated under base and extended predictor frameworks. Models were assessed on a held-out test set for discrimination, threshold-dependent performance, calibration and interpretability. The prevalence of overweight/obesity was 31.8%. Male sex and Black or African American and Hispanic/Latino race/ethnicity were associated with higher odds of overweight/obesity, whereas higher moderate-to-vigorous physical activity and more frequent breakfast intake were associated with lower odds. Extended models outperformed their corresponding base models: ROC-AUC increased from 0.6241 to 0.6894 for Random Forest and from 0.6217 to 0.6724 for XGBoost. RF-Extended was selected for calibration and interpretation. Isotonic recalibration preserved discrimination while improving probability reliability, reducing the Brier score from 0.2126 to 0.1960 and expected calibration error from 0.1160 to 0.0213. Exploratory SHAP analysis identified sex, race/ethnicity, physical activity, sleep duration, age and breakfast frequency among the features contributing most strongly to RF-Extended output. Because predictors and BMI status were measured during the same survey cycle, the models classify current status and should not be interpreted as predicting future obesity. Any screening or prevention application, including potential use in Vietnamese adolescent populations, would require prospective and external validation.