Abstract / Summary
Osteosarcoma is an aggressive bone malignancy in children and adolescents, where early and accurate classification is essential for improving clinical outcomes. While deep learning methods have shown promising results in medical image analysis, most existing approaches rely on single-modality data or assume patient-level correspondence in multimodal learning, which is rarely available in real-world clinical settings. In addition, many fusion strategies are designed for paired datasets and depend on fixed weighting schemes, limiting their applicability to heterogeneous and independently collected data. To address these challenges, this study proposes a modality-agnostic ensemble learning framework across independent cohorts for osteosarcoma classification using histopathology and X-ray images. Each modality is first modeled independently using EfficientNet-B0, and decision-level integration is performed without requiring cross-modal alignment. Two fusion strategies are investigated: fixed-weight probability fusion and a proposed adaptive sample-level gating network that dynamically learns fusion weights based on modality-specific predictive confidence. Experiments conducted on publicly available histopathology (TCIA) and X-ray (BTXRD) datasets demonstrate that the proposed adaptive fusion approach achieves superior performance compared to fixed fusion. The adaptive gating network reaches an accuracy of 97.08%, outperforming fixed-weight fusion (94.15%) and unimodal baselines. These results confirm that adaptive decision-level fusion can effectively leverage complementary information across unpaired datasets. Overall, the proposed framework provides a practical solution for multimodal medical image classification under data fragmentation conditions, where paired datasets are not available.