Abstract / Summary
Abstract Accurate segmentation of lesion regions in skin images plays a critical role in computer-aided diagnosis systems for skin lesions. However, due to the inherently complex structure of skin lesion images and the low contrast between some lesion regions and the surrounding background, achieving high-precision segmentation remains a significant challenge. Traditional convolutional neural networks struggle to capture long-range contextual dependencies, whereas Transformers excel at modeling such global relationships, offering a highly promising solution to this problem. To effectively combine the strengths of both network architectures, this paper integrates the spatial localization precision of convolutional neural networks with the semantic understanding capability of Transformers, and designs a dedicated feature interaction module to achieve deep complementary fusion of their advantages. Furthermore, considering the substantial variations in the size and morphology of skin lesions, this paper introduces a lightweight scale-aware attention mechanism that enhances multi-scale feature extraction capabilities while maintaining network efficiency. Extensive comparative experiments are conducted on two public datasets, ISIC2018 and HAM10000. The experimental results demonstrate that the proposed FIMA-Net model achieves Dice coefficients of 88.41% on the ISIC2018 dataset and 94.61% on the HAM10000 dataset, outperforming existing state-of-the-art algorithms in both quantitative metrics and visual segmentation quality. In particular, it shows significantly improved robustness when handling challenging cases, while achieving a favorable balance between segmentation accuracy and computational cost.