Infrared Small Target Detection Enhanced by Multi-Dimensional Fusion Attention
-
摘要: 针对红外小目标在深层卷积神经网络中特征弥散、易受复杂背景干扰而导致检测性能下降的问题,该文提出一种基于多维融合注意力机制的红外小目标检测方法。设计通道-空间注意力特征增强模块,通过并行分支结构聚合跨维度显著性特征,抑制目标特征在深层网络的弥散趋势;通过自适应跨维度信息交互,建立通道-空间的跨域依赖关系,增强微弱目标在深层网络的特征响应。该方法仅需极低计算复杂度即可灵活嵌入基线模型,实现目标深度特征聚拢。在NUDT-SIRST数据集上的实验表明,所提方法在高杂波、低信噪比场景下表现出更强的鲁棒性和泛化能力。基于自研的轻量化智能处理单元进行边缘部署性能验证,实测单帧推理时延为46.7 ms,满足工程应用中实时性与高精度要求。Abstract:
Objective Infrared imaging offers advantages including long operating range, wide coverage, high concealment, and all-weather operation, making it suitable for aerospace surveillance, maritime emergency rescue, forest fire monitoring, and remote sensing. Infrared sensors mounted on satellites and aircraft typically acquire long-range images in which targets occupy only a few pixels and lack discriminative texture and shape features. Moreover, weak thermal radiation signals are easily overwhelmed by background clutter. Existing infrared small target detection networks face two major challenges. First, repeated downsampling used to enlarge the receptive field causes small-target features to disappear in deep networks. Second, small-target features become increasingly diffused during deep feature extraction. Therefore, achieving accurate and reliable infrared small target detection under long-range imaging and complex background conditions remains an active research topic. Methods A Multi-Dimensional Fusion Attention Module (MFAM) based on channel-spatial attention is proposed to address feature diffusion caused by the small size of infrared targets. The proposed module captures feature dependencies across the channel, height, and width dimensions. Feature fusion and cross-dimensional interaction are jointly exploited to suppress the diffusion of small-target features in deep networks and strengthen the representation of weak infrared targets. Channel attention and spatial attention are applied in parallel directly to the input feature map, enabling feature extraction and information interaction in the channel and spatial domains, respectively. In the channel domain, compressed features are processed using a Multi-Layer Perceptron (MLP) to model inter-channel dependencies. In the spatial domain, Global Average Pooling (GAP) and Global Maximum Pooling (GMP) encode spatial information along the height and width dimensions. The outputs of the two branches are fused through a Sigmoid activation function to generate the final attention map. Unlike conventional hybrid attention mechanisms that connect channel attention and spatial attention sequentially, the proposed parallel architecture performs refined encoding of the channel, height, and width dimensions directly from the original input feature map. This design preserves both shallow spatial details and deep contextual semantics while reducing information loss during feature propagation. Owing to its plug-and-play design, MFAM can be seamlessly integrated into backbone networks such as ResNet and DNA-Net without introducing complex additional structures, demonstrating excellent compatibility. Results and Discussions The proposed method is evaluated on the publicly available NUDT-SIRST dataset. After MFAM is integrated into the baseline DNA-Net, the Intersection over Union (IoU), detection rate (Pd), and false alarm rate (Fa) reach 87.34%, 98.72%, and 3.22×10–6, respectively. Compared with CBAM, BAM, GAM, CA, and TA, MFAM improves IoU by 0.40%, 1.68%, 2.46%, 2.11%, and 1.61%, respectively (Table 1). CBAM applies spatial attention after channel attention, which increases the risk of information loss during feature propagation. CA and BAM rely solely on GAP for feature encoding and therefore fail to exploit the complementary information provided by GMP. TA models cross-dimensional dependencies through rotation operations but cannot achieve simultaneous interaction among the channel, height, and width dimensions. By jointly integrating channel and spatial attention, MFAM enables effective cross-domain information interaction. The combined use of GAP and GMP further strengthens contextual and local feature representation, resulting in more accurate infrared small target detection. MFAM is also incorporated into ALCNet and AMFU-Net, improving IoU by 0.43% and 0.28%, respectively, compared with CBAM (Table 2). Ablation experiments further demonstrate the effectiveness of the proposed design. Compared with the serial attention architecture, the parallel fusion strategy improves IoU and Pd by 0.37% and 0.39%, respectively, while reducing Fa by 1.41×10–6. These improvements result from applying channel attention and spatial attention directly to the original input feature map, thereby alleviating the diffusion of small-target features in deep networks. To verify inference performance on edge devices, the proposed method is deployed on an FPGA-GPU heterogeneous platform based on an FPGA and an NVIDIA Jetson AGX Xavier. Experimental results demonstrate successful edge deployment, with an average inference latency of 46.7 ms per 256×256 image (Fig. 8), satisfying real-time processing requirements. Conclusions A MFAM is proposed for infrared small target detection. The module effectively aggregates channel and spatial salient features while enabling adaptive cross-dimensional interaction, thereby improving the preservation of small-target features in deep networks and enhancing detection robustness. Owing to its lightweight plug-and-play design, MFAM can be readily integrated into existing detection networks. Experimental results demonstrate superior performance over existing methods in terms of IoU, Pd, and Fa. Furthermore, a lightweight intelligent processing unit based on an FPGA-GPU heterogeneous platform is developed to enable real-time deployment of the proposed algorithm on edge devices, demonstrating its practicality for engineering applications. -
Key words:
- Infrared small target /
- Attention mechanism /
- Feature fusion /
- Edge inference
-
表 1 不同算法目标检测性能(NUDT-SIRST)
类型 算法 IoU(%) Pd(%) Fa(×1e–6) 传统算法 Top-Hat 20.70 78.41 166.70 IPI 17.76 74.49 41.23 RIPT 29.44 91.85 344.30 深度学习
算法DNA-Net(基线) 85.39 97.67 9.49 +SE 86.20 98.64 3.98 +SK 85.30 98.41 12.80 +ECA 85.50 98.41 6.78 +CBAM 86.94 98.20 3.24 +BAM 85.66 98.20 5.79 +GAM 84.88 98.41 33.96 +CA 85.25 97.57 9.77 +TA 85.75 98.31 12.82 +MFAM(本文) 87.34 98.72 3.22 表 2 注意力模块嵌入不同基线模型的性能
算法 IoU(%) Pd(%) Fa(×10–6) ALC-Net (基线) 71.90 97.14 29.00 +CBAM 72.15 95.24 4.16 +MFAM (本文) 72.58 97.46 8.92 AMFU-Net (基线) 81.17 97.04 32.24 +CBAM 83.04 98.41 15.49 +MFAM (本文) 83.32 97.35 11.40 表 3 不同注意力融合方式的性能对比
算法 IoU(%) Pd(%) Fa(×1e–6) 基线模型 85.39 97.67 9.49 +C 86.78 98.18 8.95 +S 85.27 97.43 13.24 +CS,串联 86.97 98.33 4.63 +CS,并联(本文) 87.34 98.72 3.22 表 4 不同信息编码方式的性能对比
算法 IoU(%) Pd(%) Fa(×10–6) 基线模型 85.39 97.67 9.49 GAP 85.86 98.12 9.43 GMP 86.98 98.43 6.75 GAP+GMP(本文) 87.34 98.72 3.22 -
[1] 张梅, 金叶, 朱金辉, 等. 特征级语义感知引导的多模态图像融合算法[J]. 电子与信息学报, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.ZHANG Mei, JIN Ye, ZHU Jinhui,et al. FSG: Feature-level Semantic-aware Guidance for multi-modal Image Fusion Algorithm[J]. Journal of Electronics Information Technology, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042. [2] LIU J, LIU Y G, ZHANG Y Q, et al. Robust infrared small target detection using a novel four-leaf model[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 1–15. doi: 10.1109/TGRS.2023.3241234. [3] 王密, 郭贝贝, 皮英冬, 等. 珞珈三号01星在轨处理技术及验证[J]. 测绘学报, 2024, 53(4): 599–609. doi: 10.11947/j.AGCS.2024.20230412.WANG Mi, GUO Beibei, PI Yingdong, et al. On-orbit processing technology and verification of Luojia-3 01 satellite[J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(4): 599–609. doi: 10.11947/j.AGCS.2024.20230412. [4] 李召良, 唐伯惠, 吴骅, 等. 热红外遥感发展历程与展望[J]. 遥感学报, 2025, 29(6): 1529–1550. doi: 10.11834/jrs.20252301.LI Zhaoliang, TANG Bohui, WU Hua, et al. Development and prospects of thermal infrared remote sensing[J]. National Remote Sensing Bulletin, 2025, 29(6): 1529–1550. doi: 10.11834/jrs.20252301. [5] HARALICK R M, STERNBERG S R, and ZHUANG Xinhua. Image analysis using mathematical morphology[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1987, PAMI-9(4): 532–550. doi: 10.1109/TPAMI.1987.4767941. [6] HAO Congyu, LI Zhengzhou, ZHANG Yuting, et al. Infrared small target detection based on adaptive size estimation by multidirectional gradient filter[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007915. doi: 10.1109/TGRS.2024.3502421. [7] 张晶晶, 曹思华, 崔文楠, 等. 基于改进顶帽变换的红外弱小目标检测[J]. 电子与信息学报, 2024, 46(1): 267–276. doi: 10.11999/JEIT230892.ZHANG Jingjing, CAO Sihua, CUI Wennan, et al. Improved top-hat transform-based algorithm for infrared dim and small target detection[J]. Journal of Electronics & Information Technology, 2024, 46(1): 267–276. doi: 10.11999/JEIT230892. [8] GAO Chenqiang, MENG Deyu, YANG Yi, et al. Infrared patch-image model for small target detection in a single image[J]. IEEE Transactions on Image Processing, 2013, 22(12): 4996–5009. doi: 10.1109/TIP.2013.2281420. [9] 李守昌, 吴滢跃. 复杂背景下双邻域局部权重对比度红外小目标检测算法[J]. 半导体光电, 2024, 45(5): 853–860. doi: 10.16818/j.issn1001-5868.2024040301.LI Shouchang and WU Yingyue. Dual-neighborhood local weighted contrast algorithm for infrared small target detection within complex backgrounds[J]. Semiconductor Optoelectronics, 2024, 45(5): 853–860. doi: 10.16818/j.issn1001-5868.2024040301. [10] 苟士淼, 刘兆瑜, 马鹏阁, 等. 基于IHBF的加权局部对比度红外小目标检测[J]. 电光与控制, 2025, 32(8): 86–91. doi: 10.3969/j.issn.1671-637X.2025.08.014.GOU Shimiao, LIU Zhaoyu, MA Pengge, et al. Weighted local contrast infrared small target detection based on IHBF[J]. Electronics Optics & Control, 2025, 32(8): 86–91. doi: 10.3969/j.issn.1671-637X.2025.08.014. [11] DAI Yimian and WU Yiquan. Reweighted infrared patch-tensor model with both nonlocal and local priors for single-frame small target detection[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2017, 10(8): 3752–3767. doi: 10.1109/JSTARS.2017.2700023. [12] 蹇渊, 黄自力, 王询. 基于随机化张量算法的红外弱小目标检测[J]. 激光技术, 2024, 48(1): 127–134. doi: 10.7510/jgjs.issn.1001-3806.2024.01.020.JIAN Yuan, HUANG Zili, and WANG Xun. Infrared small target detection based on randomized tensor algorithm[J]. Laser Technology, 2024, 48(1): 127–134. doi: 10.7510/jgjs.issn.1001-3806.2024.01.020. [13] SUN Yang, YANG Jungang, LONG Yunli, et al. Infrared small target detection via spatial-temporal total variation regularization and weighted tensor nuclear norm[J]. IEEE Access, 2019, 7: 56667–56682. doi: 10.1109/ACCESS.2019.2914281. [14] LIU Ting, YANG Jungang, LI Boyang, et al. Nonconvex tensor low-rank approximation for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60: 5614718. doi: 10.1109/TGRS.2021.3130310. [15] 胡亮, 杨德贵, 赵党军, 等. 改进非凸估计与非对称时空正则化的红外小目标检测方法[J]. 国防科技大学学报, 2024, 46(3): 180–194. doi: 10.11887/j.cn.202403018.HU Liang, YANG Degui, ZHAO Dangjun, et al. Infrared small target detection method based on improved non-convex estimation and asymmetric spatial-temporal regularization[J]. Journal of National University of Defense Technology, 2024, 46(3): 180–194. doi: 10.11887/j.cn.202403018. [16] LI Fenghong, RAO Peng, SUN Wen, et al. A new motion feature-enhanced multiframe spatial-temporal infrared target detection network[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5006819. doi: 10.1109/TGRS.2025.3603784. [17] XU Jielei, HAN Xinheng, WANG Jiacheng, et al. PCLC-Net: Parallel connected lateral chain networks for infrared small target detection[J]. Remote Sensing, 2025, 17(12): 2072. doi: 10.3390/rs17122072. [18] 龙畅, 张弦, 刘艳阳, 等. 天基红外小目标探测技术新进展、挑战与对策[J]. 光电工程, 2025, 52(12): 250300. doi: 10.12086/oee.2025.250300.LONG Chang, ZHANG Xian, LIU Yanyang, et al. Recent progress, challenges, and countermeasures in the detection of dim small targets using space-based infrared systems[J]. Opto-Electronic Engineering, 2025, 52(12): 250300. doi: 10.12086/oee.2025.250300. [19] LI Fenghong, RAO Peng, SUN Wen, et al. A low signal-to-noise ratio infrared small-target detection network[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 8643–8658. doi: 10.1109/JSTARS.2025.3550581. [20] LIU Ming, DU Haoyuan, ZHAO Yuejin, et al. Image small target detection based on deep learning with SNR controlled sample generation[J]. Current Trends in Computer Science and Mechanical Automation, 2017, 1(1): 211–220. doi: 10.1515/9783110584974-025. [21] HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition[C]. The 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, USA, 2016: 770–778. doi: 10.1109/CVPR.2016.90. [22] 盛卫东, 吴双林, 肖超, 等. 可微稀疏掩模引导的红外小目标快速检测网络[J]. 电子与信息学报, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989.SHENG Weidong, WU Shuanglin, XIAO Chao, et al. Differentiable sparse mask guided infrared small target fast detection network[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989. [23] ZHANG Mingjin, ZHANG Rui, YANG Yuxiang, et al. ISNet: Shape matters for infrared small target detection[C]. The 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 867–876. doi: 10.1109/CVPR52688.2022.00095. [24] ZHANG Tianfang, LI Lei, CAO Siying, et al. Attention-guided pyramid context networks for detecting infrared small target under complex background[J]. IEEE Transactions on Aerospace and Electronic Systems, 2023, 59(4): 4250–4261. doi: 10.1109/TAES.2023.3238703. [25] LI Chengyu, ZHANG Yan, SHI Zhiguang, et al. Moderately dense adaptive feature fusion network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5616712. doi: 10.1109/TGRS.2024.3381006. [26] LIN Fanzhao, BAO Kexin, LI Yong, et al. Learning contrast-enhanced shape-biased representations for infrared small target detection[J]. IEEE Transactions on Image Processing, 2024, 33: 3047–3058. doi: 10.1109/TIP.2024.3391011. [27] ZHOU Zongwei, RAHMAN SIDDIQUEE M M, TAJBAKHSH N, et al. Unet++: A nested u-net architecture for medical image segmentation[C]. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support 4th International Workshop, Granada, Spain, 2018: 3–11. doi: 10.1007/978-3-030-00889-5_1 . [28] DAI Yimian, WU Yiquan, ZHOU Fei, et al. Asymmetric contextual modulation for infrared small target detection[C]. The 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, USA, 2021: 949–958. doi: 10.1109/WACV48630.2021.00099 . [29] DAI Yimian, WU Yiquan, ZHOU Fei, et al. Attentional local contrast networks for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 59(11): 9813–9824. doi: 10.1109/TGRS.2020.3044958. [30] LI Boyang, XIAO Chao, WANG Longguang, et al. Dense nested attention network for infrared small target detection[J]. IEEE Transactions on Image Processing, 2023, 32: 1745–1758. doi: 10.1109/TIP.2022.3199107. [31] LIN T Y, DOLLÁR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]. The 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, USA, 2017: 2117–2125. doi: 10.1109/CVPR.2017.106. [32] 王媛彬, 吴冰超. 基于自适应特征融合和注意力机制的变电设备红外图像识别[J]. 电子与信息学报, 2024, 46(9): 3749–3756. doi: 10.11999/JEIT231047.WANG Yuanbin and WU Bingchao. Infrared Image Recognition of Substation Equipment Based on Adaptive Feature Fusion and Attention Mechanism[J]. Journal of Electronics Information Technology, 2024, 46(9): 3749–3756. doi: 10.11999/JEIT231047. [33] GUO Huinan, ZHANG Nengshuang, ZHANG Jing, et al. Location-guided dense nested attention network for infrared small target detection[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17: 18535–18548. doi: 10.1109/JSTARS.2024.3472041. [34] HU Jie, SHEN Li, and SUN Gang. Squeeze-and-excitation networks[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 7132–7141. doi: 10.1109/CVPR.2018.00745. [35] LI Xiang, WANG Wenhai, HU Xiaolin, et al. Selective kernel networks[C]. The 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 510–519. doi: 10.1109/CVPR.2019.00060. [36] WANG Qilong, WU Banggu, ZHU Pengfei, et al. ECA-Net: Efficient channel attention for deep convolutional neural networks[C]. The 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2020: 11531–11539. doi: 10.1109/CVPR42600.2020.01155. [37] WOO S H, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. The 15th European Conference on Computer Vision-ECCV 2018, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1. [38] PARK J, WOO S, LEE J Y, et al. BAM: Bottleneck attention module[C]. The 2018 British Machine Vision Conference (BMVC), Newcastle, UK, 2018: 1–20. doi: 10.5244/C.32.1. [39] LIU Yichao, SHAO Zongru, and HOFFMANN N. Global attention mechanism: Retain information to enhance channel-spatial interactions[EB/OL]. https://arxiv.org/abs/2112.05561, 2021. [40] ZHONG Yuhan, SHI Zhiguang, ZHANG Yan, et al. CSAN-UNET: Channel spatial attention nested UNet for infrared small target detection[J]. Remote Sensing, 2024, 16(11): 1894. doi: 10.3390/rs16111894. [41] ZUO Zhen, TONG Xiaozhong, WEI Junyu, et al. AFFPN: Attention fusion feature pyramid network for small infrared target detection[J]. Remote Sensing, 2022, 14(14): 3412. doi: 10.3390/rs14143412. [42] HOU Qibin, ZHOU Daquan, and FENG Jiashi. Coordinate attention for efficient mobile network design[C]. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, USA, 2021: 13708–13717. doi: 10.1109/CVPR46437.2021.01350. [43] MISRA D, NALAMADA T, ARASANIPALAI A U, et al. Rotate to attend: Convolutional triplet attention module[C]. 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, USA, 2021: 3138–3147. doi: 10.1109/WACV48630.2021.00318. -
下载: