Abstract / Summary
Background: Albuminuria is central to staging and risk prediction in chronic kidney disease (CKD), yet urine albumin to creatinine ratio (UACR) testing is performed far less often than serum creatinine testing. Whether routinely available clinical variables can identify patients likely to have albuminuria, and so substitute for or prioritize testing, is unclear. Methods: We analyzed adults in NHANES 2015 to 2018 with an eGFR of 30 to 59 mL/min/1.73m2 (CKDEPI 2021). The outcome was UACR >=30 mg/g. Seventeen features derived from demographics, serum creatinine, eGFR, blood pressure, diabetes status and self-reported kidney disease were used; UACR was excluded. Logistic regression, random forest and gradient boosting were evaluated with 10 times repeated stratified 5 fold cross validation. We report AUC with bootstrap 95% CIs, DeLong tests, calibration, survey weighted AUC, diabetes subgroups and a rule-out analysis. Results: Of 687 participants, 239 (34.8%) had UACR >=30 mg/g. Discrimination was modest and similar across models: logistic regression AUC 0.723 (95% CI 0.684 to 0.762), random forest 0.721 (0.680 to 0.761; DeLong p = 0.84 vs logistic regression) and gradient boosting 0.719 (0.679 to 0.759; p = 0.76). Logistic regression and random forest were reasonably calibrated (slopes 0.87 and 0.98); gradient boosting was overfitted (slope 0.64). Discrimination was lower in participants without diabetes (logistic regression AUC 0.636 vs 0.722 with diabetes). Using logistic regression to rule out testing at 90% sensitivity would have spared 22.7% of participants from testing but missed 23 of 239 cases; at 95% sensitivity, only 16.3% would have been spared. Conclusions: Routine clinical variables identify albuminuria only modestly, and flexible machine-learning models offer no advantage over logistic regression. No model allowed a meaningful share of patients to forgo UACR testing without missing cases, supporting universal UACR testing in moderate CKD. Keywords: chronic kidney disease; albuminuria; urine albumin to creatinine ratio; machine learning; logistic regression; NHANES; screening