Abstract / Summary
In hypertension screening, previously diagnosed cases provide reliable positive labels, whereas individuals without a prior diagnosis comprise both undiagnosed hypertensive and non-hypertensive individuals, creating a canonical positive-unlabeled (PU) learning setting. We propose PU-Boost, a two-stage reliable-sample reconstruction framework that uses only readily obtainable demographic, behavioral, medical-history, and anthropometric variables. During training, only previously diagnosed hypertension cases are treated as labeled positives (P), while all remaining individuals are treated as unlabeled (U); no additional reference labels are supplied to the two-stage reconstruction procedure. Elastic Net first provides a global ranking of P versus U, after which a random forest refines the uncertain region and highly ambiguous observations are rejected. XGBoost is then trained on the reconstructed class-balanced pseudo-labeled sample. In the independent test set, PU-Boost achieved a sensitivity of 0.817 and a balanced accuracy of 0.682, the highest observed values among the compared models. It produced 729 false negatives, 41.1% fewer than standard XGBoost, although specificity decreased to 0.547. SHAP analysis indicated that established hypertension-related features, including age, body mass index, waist circumference, and family history, contributed strongly to model output. PU-Boost therefore provides a statistically motivated framework for reducing missed cases when screening for undiagnosed hypertension under incomplete diagnostic labeling.