Abstract / Summary
Abstract Background Rising global obesity imposes substantial public health burdens worldwide. Traditional standalone machine learning models lack robust clustering frameworks, restricting their performance in multi-grade obesity classification tasks. Identifying and ranking key risk factors is essential to realize precise population level obesity prevention and targeted intervention. This study aims to propose a multi-level obesity hybrid clustering classification model to obtain the importance ranking of obesity-related risk factors and provide scientifically supported risk decision-making. Methods This study used a publicly available obesity level dataset containing 20,758 valid samples for data cleaning, feature encoding, and Pearson correlation analysis. A hybrid model combining K-means clustering and extreme gradient boosting (XGBoost) was constructed. Different learning rate settings were compared, and 10-fold cross-validation was employed to avoid over-fitting. The performance of this model was compared with K-means integration baseline models such as support vector machines, decision tree, random forest, and multilayer perceptron. The proposed algorithm K-means + XGBoost was used to calculate feature importance and to rank and analyze obesity key lifestyle risk factors. Results The optimized K-means + XGBoost hybrid model with a learning rate of 0.1 yielded 98.28% training accuracy and 90.15% test accuracy. Based on the K-means + XGBoost scores, the risk factors ranking for the multi-grade obesity level dataset is Gender (0.3110) > Weight (0.1719) > FCVC (vegetable intake frequency, 0.0375) > FAVC (high calorie food intake frequency, 0.0312) > Height (0.0303) > SCC (sugary drink consumption, 0.0263) > CALC (alcohol intake frequency, 0.0238) > Age (0.0185) > CAEC (between meal snacking frequency, 0.0179) > NCP (daily main meal number, 0.0169) > CH2O (daily water intake, 0.0169) > SMOKE (smoking status, 0.0168) > family history with overweight (0.0155) > MTRANS (main transportation mode, 0.0148) > TUE (technological device usage time, 0.0100) > FAF (physical activity frequency, 0.0090). Conclusions The proposed K-means + XGBoost hybrid framework achieves favorable classification performance for multi-grade obesity level dataset. The ranked risk factors provide practical evidence for conducting targeted public health screening and implementing differentiated obesity intervention measures.