Abstract / Summary
The precise and prompt identification of ocular conditions, such as diabetic retinopathy (DR), diabetic macular edema (DME), or glaucoma, often depends on the collaborative analysis of medical imaging techniques such as fundus photography or optical coherence tomography (OCT), alongside the patient’s demographic and acquisition metadata. This research introduced a unique unimodal and multimodal neural network (multimodal net) architecture that leverages complementary information for disease categorization. The unimodal approach has an individual architecture (an image branch and a tabular branch), whereas the multimodal net has a dual-branch architecture (an image branch that leverages a pre-trained ResNet-18 to extract deep visual features from OCT images and a tabular branch that includes a multi-layer perceptron [MLP] to analyze demographic and acquisition metadata factors). The features obtained from both modalities are concatenated and then integrated through a combined branch to predict the target illness class. By combining these diverse data types, the model aimed to gain a holistic understanding of the patient’s condition, potentially yielding a more robust and demographic-relevant diagnostic system than those that rely on a single modality. The trial findings showed that the multimodal method, which combines patient demographics and image details with retinal OCT images, performs slightly better than unimodal approaches. With 94.92% accuracy and an Misclassification Rate (MCC) of 0.0508, the EfficientNet-B0 + MLP configuration significantly outperformed the ResNet18 baseline (86.9% accuracy) for the image-level split and achieved 0.924 validation accuracy with 0.871 MCC for the patient-level split. Further, we validated the fusion layer’s ability to capture synergies between visual and tabular data through five-fold cross-validation, which improved performance to an overall accuracy with a mean and standard deviation of 0.932 ± 0.006 and macro-F1 of 0.876 ± 0.023 for patient-level splits. Moreover, we compared our proposed model with image-only, tabular-only, and gated multimodal fusion (GMU)-fused structures. The image-only model showed strong predictive performance (91.23% accuracy, 0.94 Area Under the Curve (AUC)) compared with the tabular-only model (64.66% accuracy), although integrating both modalities yielded even better outcomes. The suggested concat (image + tabular) framework performed best, achieving 92.43% accuracy (95% CI: 0.895–0.947), a macro-AUC of 0.9703, and a minimal misclassification rate of 0.77, illustrating the benefits of combining demographic and acquisition tabular attributes with visual information. We adopted EfficientNet-B0 as our primary encoder. To assess whether performance depends on this choice, we also evaluated ResNet-18, ViT, Swin, MaxViT, and CoAtNet. All encoders achieved comparable macro-F1 (0.86–0.89), with differences within the range of single-split variation, confirming that the results are not specific to a single architecture. Grad-CAM (Gradient-weighted Class Activation Mapping) visualizations were also reported.