Abstract / Summary
Multimodal retinal image registration remains challenging due to nonlinear intensity discrepancies and insufficient feature representation across modalities. To address these limitations, we propose VesselRegNet, a joint retinal vessel segmentation and registration framework based on multi-scale feature fusion. The model integrates a multi-scale semantic feature extraction module with a shared prediction head and rotation-equivariant constraints to enhance feature consistency. Unlike conventional approaches relying on binary segmentation masks, the proposed method utilizes dense semantic features to directly guide deformable registration. Experimental results demonstrate that VesselRegNet achieves an average Dice coefficient of 0.664 and SSIM of 0.645, achieving a relative improvement of 4.1% in Dice and 2.5% in SSIM over the strongest baseline (RetinaRegNet), and up to 11.9% improvement over general unsupervised registration frameworks. The proposed multi-scale fusion and joint learning strategy improves registration accuracy under complex retinal deformations.