Abstract / Summary
This is the revised second version of the research article, “Stroke Risk Prediction Using Logistic Regression and Classification Trees in R.” This version updates the statistical analysis, results, interpretation, and reproducibility materials of the original publication. The revised analysis was independently reproduced using the same dataset and the original R workflow, and inconsistencies identified during reproduction were corrected. The revised version provides updated logistic regression and classification-tree results, including appropriate evaluation of class imbalance, sensitivity, specificity, ROC-AUC, threshold-dependent performance, model selection criteria, and limitations. The Appendix has also been updated to include the code used to reproduce the reported analyses and performance measures. The revised findings show that raw accuracy is strongly affected by the low prevalence of stroke cases in the dataset. The final logistic regression model achieved 95.23% raw test accuracy but 0% sensitivity at the standard 0.5 threshold, while its ROC-AUC was 0.833. The classification tree achieved 95.31% test accuracy with 6.6% sensitivity at its coded 0.3 threshold. These results are interpreted as evidence that raw accuracy alone is insufficient for evaluating predictive performance on this imbalanced dataset. This version supersedes the original version of the article while preserving the publication history of the work.