Abstract / Summary
Emergency Department (ED) length of stay (LOS) is a critical performance indicator in healthcare systems, influencing patient outcomes, overcrowding, and resource utilization. Predicting LOS can enhance patient flow and resource management. While traditional statistical methods have been used, the advent of machine learning (ML) has introduced more sophisticated analytical approaches, though their comparative advantage over traditional methods remains context-dependent. A systematic search was conducted across PubMed, Embase, Web of Science, and Scopus using terms related to ED LOS and predictive modeling. Eligible studies were English-language studies published from January 2014 through September 2024 that developed or validated an ED LOS prediction model using a traditional statistical, machine-learning, or deep-learning approach. Studies were excluded if they did not develop or validate an eligible ED LOS prediction model or if ED LOS was not the prediction target. Two reviewers independently screened all titles and abstracts, and three reviewers independently assessed the full texts of potentially eligible reports. Two reviewers independently extracted data from all included studies, with disagreements adjudicated by an emergency physician who served as the study supervisor. Of 4,773 records identified, 18 met the inclusion criteria. The studies spanned multiple continents, with the majority conducted post-2020. Decision-tree models were evaluated most frequently, followed by random forests, logistic regression, and support vector machines or support vector regression. The area under the receiver operating characteristic curve (AUROC) and accuracy were the most common performance metrics; however, direct cross-study statistical comparisons were precluded by substantial methodological heterogeneity (e.g., varied LOS cut-offs, patient populations, and validation methods). Extracted predictors commonly represented laboratory testing and ED operational factors, although inconsistent feature-importance reporting precluded comparison of their predictive importance. Calibration was rarely reported, with only two studies providing calibration assessments. Reported performance varied across modeling approaches, predictor sets, LOS outcome definitions, validation procedures, and ED settings. Because substantial methodological heterogeneity limits direct performance comparisons, future research should adopt standardized reporting frameworks (e.g., TRIPOD) and explore the impact of different LOS cut-off values to facilitate meaningful evidence synthesis and improve resource allocation. Not applicable.