Abstract / Summary
Background: Epigenetic deregulation is a defining feature of leukemia, and DNA methylation and chromatin accessibility profiles encode the cell of origin, lineage and clinical behaviour of the disease. Artificial intelligence (AI) and machine learning (ML) can extract predictive patterns from these high-dimensional data, but the published models have not been systematically appraised. We systematically reviewed AI models built on epigenomic data in leukemia for diagnosis, prognosis, relapse and treatment-response prediction, and measurable residual disease (MRD) detection.
Methods: This review followed PRISMA 2020 and the SWiM guideline. PubMed was searched on 7 October 2026 with three structured queries, supplemented by a broader PubMed query, citation searching and web searching. Eligible studies were peer-reviewed primary studies that applied an ML or deep-learning algorithm to DNA methylation, chromatin-accessibility or histone-modification data from human leukemia samples and reported a quantitative diagnostic, prognostic, response or MRD outcome. Risk of bias and applicability were assessed with PROBAST, informed by PROBAST+AI. Owing to heterogeneity, results were synthesized narratively and in structured tables, without meta-analysis.
Results: Of 68 unique database records and 32 reports from other sources, 30 studies (2010-2026; 22 since 2020) were included: 16 in acute myeloid leukemia, 4 in acute lymphoblastic leukemia, 2 spanning acute leukemia lineages, 3 in juvenile myelomonocytic leukemia and 5 in chronic lymphocytic leukemia; 10 were pediatric. DNA methylation was the input in 28 studies and chromatin accessibility in 2; no eligible study modelled histone marks. Tree ensembles (n = 9) and support vector machines (n = 5) were the most frequent algorithms. Diagnostic classifiers reported accuracy, concordance or F1 scores of 0.87-1.00, and two classifiers were applied to nanopore sequencing data, one returning acute leukemia subtype calls within about 2 hours. Prognostic epigenomic scores were associated with survival (reported hazard ratios 2.36-10.82), but discrimination fell on transport to independent cohorts (C-index 0.677 internally vs 0.529 in one external cohort). Fifteen studies validated models on an external epigenomic cohort, one prospectively. Overall risk of bias was high in 21 studies, unclear in 8 and low in 1, driven by the analysis domain (small samples, absent calibration, possible data leakage). Calibration was reported in 3 studies and code was available for 11.
Conclusions: Epigenomic AI models classify leukemia subtypes with high reported accuracy and can deliver same-day diagnoses, and methylation-based scores add prognostic information. However, the evidence is dominated by retrospective, high-risk-of-bias development studies with limited external validation and almost no prospective or impact evaluation. Prospective, multi-center validation with calibrated, transparently reported models is required before clinical adoption.