Abstract / Summary
Substantial variability in organoid differentiation into well-defined target cell populations necessitates objective, nondestructive quality assessment early during differentiation to support successful organoid-based therapies. We propose a triplet-trained Vision Transformer (Triplet-ViT) to predict differentiation outcomes of hypothalamic–pituitary organoids using a publicly available dataset comprising 1,500 bright-field images and compares its performance with EfficientNetV2-S and their 50:50 ensemble. Triplet-ViT achieved 71.33% accuracy, which exceeds the previously reported ViT benchmark by 5.63 percentage points. EfficientNetV2-S achieved the highest overall accuracy (73.67%) and recall (88.00%) for Category C organoids, which represented poorly differentiated organoids that were considered candidates for early exclusion from further culture, whereas Triplet-ViT achieved the highest Category C precision (86.81%) and specificity (94.00%) among the evaluated models and their ensemble. To assess triplet learning, we compared Triplet-ViT with a CE-only ViT under matched conditions using a five-seed, five-fold cross-validation within a 1,200-image training set. The accuracy and macro-F1 remained similar, whereas triplet learning increased C recall and C F1 by 2.70 and 1.02 percentage points, respectively. Advanced image-based deep learning, including triplet learning, may enable a more accurate, nondestructive, and standardized prediction of organoid differentiation, although validation across diverse cell lines and organoid systems is required before broader clinical application.