Abstract / Summary
Background/Objectives: Multimodal artificial intelligence (AI) is increasingly used in radiology to combine medical images with radiology reports, clinical narratives, structured health records, laboratory measurements, and other patient data. These systems may support more context-aware interpretation, reporting, and clinical decision-making than unimodal approaches. This narrative review examines recent technical and clinical advances in multimodal radiology AI and identifies barriers to responsible implementation. Methods: Relevant biomedical and technical literature was identified through targeted searches of PubMed, IEEE Xplore, ACM Digital Library, Web of Science, Scopus, arXiv, ScienceDirect, SpringerLink, and Google Scholar. A documented screening process identified 111 reviewed publications from 1154 records. Study selection and data extraction were conducted by one reviewer without independent verification. Original research, reviews, and commentaries, including selected preprints, were synthesized thematically. Results: The field has progressed from early and late feature fusion toward cross-modal attention, contrastive image–text pretraining, vision–language models, multimodal large language models, and general-purpose foundation models. Applications include diagnostic classification, prognostic modelling, image–text retrieval, visual question answering, clinical decision support, and automated radiology report generation. Despite promising technical results, comparison across studies remains difficult because of heterogeneous datasets, tasks, metrics, and validation designs. Clinical translation is further limited by scarce external and prospective validation, uncertain interpretability, hallucination and omission risks, privacy and fairness concerns, and limited real-world workflow evaluation. Conclusions: Multimodal AI may enable more clinically informed radiological interpretation and reporting, but progress in benchmark performance has outpaced evidence of safety, generalizability, and clinical utility. Future research should prioritize multicentre evaluation, clinically meaningful metrics, transparent reporting, robust safety assessment, and workflow-centred implementation.