Abstract / Summary
Abstract Predicting HIV inhibitors is important for accelerating antiviral drug discovery. This study compares fingerprint-based machine learning, graph-based deep learning, and ensemble approaches for classifying compounds in the MoleculeNet HIV dataset, which contains 41,127 SMILES-encoded molecules and is highly imbalanced (3.51% active). Morgan fingerprints (radius 2, 2048 bits) were used with Random Forest, SVM, XGBoost, CatBoost, LightGBM, and feedforward neural networks, while molecular graphs with atom and bond features were processed using a Graph Attention Network (GAT). Class imbalance was addressed through resampling, and Optuna was used for hyperparameter tuning. A hybrid ChemBERTa-GAT-fingerprint model with cross-attention was also evaluated, together with a stacking ensemble using Random Forest, LightGBM, CatBoost, and XGBoost as base learners and an FNN as meta-learner. XGBoost achieved 0.8567 accuracy and 0.8508 ROC-AUC; the GAT achieved 0.8100 accuracy and 0.8534 ROC-AUC. The stacking ensemble achieved 0.8521 accuracy and 0.8499 ROC-AUC, while the hybrid model obtained 0.8400 ROC-AUC and 0.816 PR-AUC. Attention-based visualizations highlighted atoms contributing strongly to GAT predictions. These results show the complementary value of fingerprint-based and graph-based molecular representations for HIV inhibitor classification and support interpretable AI-assisted molecular screening.