Abstract / Summary
The emergence of drug resistance creates a critical gap in antitubercular lead prioritization, as small-molecules identified against drug-susceptible (DS) strains may not retain similar potential against drug-resistant (DR) isolates. Here, we developed an AI/ML workflow that separately prioritizes small molecules against DR and DS isolates and models MIC, MIC90, and MIC99 as distinct quantitative endpoints rather than combining them into a single activity label. After curation, the binary classification datasets contained 25645 and 3749 compounds against DS and DR strains respectively. Molecular representations included RDKit descriptors and Morgan fingerprints, ChemBERTa, MolCLR, and Uni-Mol, and all comparisons were performed with scaffold-held-out cross-validation to reduce analogue leakage. The strongest DS binary classifier used RDKit/Morgan features with ExtraTrees (balanced accuracy 0.782; AUROC 0.864; AUPRC 0.870), whereas the best DR classifier used the same representation with histogram gradient boosting (balanced accuracy 0.762; AUROC 0.836; AUPRC 0.880). For regression, the best overall model for each of the six DS/DR endpoint combinations also used RDKit/Morgan features with a tree-based regressor, with pooled out-of-fold RMSE values of 0.604-0.784 pMIC. Learned molecular embeddings retained substantial predictive signal but did not provide a consistent advantage under scaffold-held-out evaluation. By separating isolate context and assay endpoint, the workflow provides a more specific basis for prioritizing compounds for testing against susceptible and resistant M. tuberculosis (M.tb). To facilitate broader use, we developed MtbAIM (https://github.com/jasdeep002/MtbAIM), an accessible prediction platform deployed through Google Colab for prioritizing compounds against DS and DR M.tb.