Abstract / Summary
Abstract The field of medical research is currently undergoing a paradigm shift as the need for automated systems to augment or offload cognitive burdens becomes increasingly apparent. Within this context, multimodal artificial intelligence (AI) offers a clear advantage over unimodal approaches by mirroring clinical practice through the integration of multiple data sources. Nonetheless, many multimodal frameworks do not process raw image data directly, relying instead on extracted information that can obscure inherent data complexities. This scoping review mapped the literature regarding multimodal AI for mental and cognitive health, specifically where one input is raw image data and the other is non-image data. Searching PubMed, PsycArticles, and IEEE Xplore, 19 articles were eligible for inclusion, from which only two presented real-world validation against clinical professionals. A salient result is the heavy reliance on publicly available data, such as that from the Alzheimer’s Disease Neuroimaging Initiative (ADNI). Consequently, research is predominantly focused on Mild Cognitive Impairment (MCI) and Alzheimer’s Disease (AD), mostly involving magnetic resonance imaging (MRI). Furthermore, most articles utilized intermediate fusion, yet there remains significant variety in fusion levels and mechanisms, ranging from simple aggregation to neural and joint learning. In conclusion, future research needs to prioritize the rigorous comparison of fusion levels and mechanisms. Additionally, the adoption of diverse, open-access databases beyond AD is essential to address a broader range of mental conditions and traits.