Infrared Small Target Detection Enhanced by Multi-dimensional Fusion Attention
-
摘要: 针对红外小目标在深层卷积神经网络中特征弥散、易受复杂背景干扰而导致检测性能下降的问题,提出一种基于多维融合注意力机制的红外小目标检测方法。设计通道-空间注意力特征增强模块,通过并行分支结构聚合跨维度显著性特征,抑制目标特征在深层网络的弥散趋势;通过自适应跨维度信息交互,建立通道-空间的跨域依赖关系,增强微弱目标在深层网络的特征响应。该方法仅需极低计算复杂度即可灵活嵌入基线模型,实现目标深度特征聚拢。在NUDT-SIRST数据集上的实验表明,所提方法在高杂波、低信噪比场景下表现出更强的鲁棒性和泛化能力。基于自研的轻量化智能处理单元进行边缘部署性能验证,实测单帧推理时延为46.7 ms,满足工程应用中实时性与高精度要求。Abstract:
Objective Infrared imaging boasts advantages such as long operating range, wide coverage, strong concealment, and all-weather visibility, making it widely applicable in aerospace target surveillance, maritime emergency rescue, forest fire monitoring, and earth remote sensing. Infrared payloads are typically mounted on platforms like satellites and aircraft, capturing images over long distances where targets appear small in the imagery and lack distinct texture and morphological features. Due to weak thermal radiation signals, targets are prone to being submerged in strong clutter. Current infrared small target detection networks face several key challenges. First, target segmentation networks expand the receptive field through continuous down-sampling, which can cause the features of infrared small targets to be easily lost in the deeper layers of neural networks. Second, features in the deeper neural network layers tend to diffuse. Therefore, under the constraints of long-range imaging and complex backgrounds, achieving high-precision and highly reliable infrared small target detection remains a research hotspot in the field of infrared imaging. Methods A channel-spatial Multi-Dimensional Fusion Attention Mechanism (MFAM) is proposed to address the challenges of small target size and feature diffusion in deep networks. Multi-dimensional features across channels, height, and width are captured, as well as the dependencies across these dimensions. By integrating feature fusion and cross-dimensional synergistic interaction, feature diffusion of small targets in deep networks is mitigated and the robustness of infrared small-target detection is enhanced. The channel attention and spatial attention modules are applied in parallel directly to the input feature maps, performing feature extraction and local information interaction along the channel and spatial dimensions, respectively. In the channel domain, features are compressed and processed through a Multi-layer Perceptron (MLP) to capture inter-channel relationships. In the spatial domain, Global Average Pooling (GAP) and Global Maximum Pooling (GMP) are employed to encode width and height dimensions. Finally, feature fusion is achieved via a Sigmoid activation function. Compared with traditional hybrid attention mechanisms, the channel attention module and spatial attention module are applied in parallel directly to the input feature maps. This approach enables refined encoding of the channel, height, and width dimensions under a global receptive field. Meanwhile, shallow features with spatial details and deep semantic features rich in contextual information are captured. The proposed MFAM features a plug-and-play design, allowing flexible integration into various baseline models such as ResNet and DNA-Net without introducing complex additional structures, which demonstrates excellent compatibility and generalization capability. Results and Discussions The publicly available dataset NUDT-SIRST is utilized to conduct a performance analysis of the proposed algorithm. By incorporating the proposed MFAM into the baseline DNA-Net, the Intersection over Union (IoU), detection rate (Pd), and false alarm rate (Fa) are 87.34%, 98.72% and 3.22×10–6, respectively. Compared to CBAM, BAM, GAM, CA, and TA, MFAM improves IoU by 0.4%, 1.68%, 2.46%, 2.11%, and 1.61%, respectively( Table.1 ). CBAM places the spatial attention module in series of the channel attention, which may lead to information loss during feature propagation. CA and BAM only employ GAP for information encoding, overlooking the role of max pooling in deep feature extraction. TA realizes cross-dimensional dependencies through rotation operations but fails to achieve simultaneous three-dimensional interaction. MFAM simultaneously integrates channel and spatial information, enabling sufficient cross-domain interaction, and combines global average and max pooling to enhance context and detailed feature extraction capabilities, thereby achieving more refined infrared small target detection. MFAM is also embedded into ALC-Net and AMFU-Net, resulting in IoU improvements of 0.43% and 0.28% compared with CBAM(Table.2 ). Through ablation experiments, detection performance is compared under different attention combination strategies. Compared to the serial fusion approach, the proposed parallel structure achieves improvements of 0.37% and 0.39% in IoU and Pd, respectively, while reducing Fa by 1.41×10–6. The advantage stems from the parallel structure applying channel and spatial attention directly to the input features, mitigating the diffusion of target features in deep networks. To validate the inference performance of the proposed algorithm on edge-side processors, a verification system is designed by using FPGA and NVIDIA Jetson AGX Xavier. Practical testing confirms that the proposed algorithm can be deployed and inferred on edge-side GPUs, with an average single-frame inference latency of 46.7 ms (Fig. 8 ) for 256×256 input images.Conclusions A Multi-dimensional Fusion Attention Module (MFAM) is constructed in this paper. The MFAM module achieves effective aggregation of channel and spatial salient features and enables cross-dimensional adaptive interaction, enhancing the preservation of target features in deep networks and ensuring robust output. The MFAM module exhibits favorable plug-and-play characteristics. Experiments demonstrate that the proposed algorithm performs better in metrics of IoU, Pd, and Fa. A lightweight intelligent processing unit based on FPGA+GPU is developed, which successfully achieves deployment of the algorithm on edge devices and enables high real-time inference, demonstrating promising engineering applicability and future application prospects. -
Key words:
- Infrared small target /
- Attention mechanism /
- Feature fusion /
- Edge intelligence
-
表 1 不同算法目标检测性能(NUDT-SIRST)
类型 算法 IoU/(%) Pd/(%) Fa/(×1e-6) 传统算法 Top-Hat 20.7 78.41 166.7 IPI 17.76 74.49 41.23 RIPT 29.44 91.85 344.3 深度学习
算法DNA-Net(基线) 85.39 97.67 9.49 +SE 86.2 98.64 3.98 +SK 85.3 98.41 12.8 +ECA 85.50 98.41 6.78 +CBAM 86.94 98.2 3.24 +BAM 85.66 98.2 5.79 +GAM 84.88 98.41 33.96 +CA 85.25 97.57 9.77 +TA 85.75 98.31 12.82 +MFAM(本文) 87.34 98.72 3.22 表 2 注意力模块嵌入不同基线模型的性能
算法 IoU/(%) Pd/(%) Fa/(×1e-6) ALC-Net (基线) 71.9 97.14 29 +CBAM 72.15 95.24 4.16 +MFAM (本文) 72.58 97.46 8.92 AMFU-Net (基线) 81.17 97.04 32.24 +CBAM 83.04 98.41 15.49 +MFAM (本文) 83.32 97.35 11.4 表 3 不同注意力融合方式的性能对比
算法 IoU/(%) Pd/(%) Fa/(×1e-6) 基线模型 85.39 97.67 9.49 +C 86.78 98.18 8.95 +S 85.27 97.43 13.24 +CS,串联 86.97 98.33 4.63 +CS,并联(本文) 87.34 98.72 3.22 表 4 不同信息编码方式的性能对比
算法 IoU/(%) Pd/(%) Fa/(×1e-6) 基线模型 85.39 97.67 9.49 GAP 85.86 98.12 9.43 GMP 86.98 98.43 6.75 GAP+GMP(本文) 87.34 98.72 3.22 -
[1] 周海, 李保权, 王怀超, 等. 低空复杂场景红外弱小目标快速精准检测[J]. 国防科技大学学报, 2023, 45(1): 74–85. doi: 10.11887/j.cn.202301008.ZHOU Hai, LI Baoquan, WANG Huaichao, et al. Fast and accurate detection of infrared dim small target in low altitude complex scenes[J]. Journal of National University of Defense Technology, 2023, 45(1): 74–85. doi: 10.11887/j.cn.202301008. [2] LIU J, LIU Y G, ZHANG Y Q, et al. Robust infrared small target detection using a novel four-leaf model[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 1–15. doi: 10.1109/TGRS.2023.3241234. [3] 王密, 郭贝贝, 皮英冬, 等. 珞珈三号01星在轨处理技术及验证[J]. 测绘学报, 2024, 53(4): 599–609. doi: 10.11947/j.AGCS.2024.20230412.WANG Mi, GUO Beibei, PI Yingdong, et al. On-orbit processing technology and verification of Luojia-3 01 satellite[J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(4): 599–609. doi: 10.11947/j.AGCS.2024.20230412. [4] 李召良, 唐伯惠, 吴骅, 等. 热红外遥感发展历程与展望[J]. 遥感学报, 2025, 29(6): 1529–1550. doi: 10.11834/jrs.20252301.LI Zhaoliang, TANG Bohui, WU Hua, et al. Development and prospects of thermal infrared remote sensing[J]. National Remote Sensing Bulletin, 2025, 29(6): 1529–1550. doi: 10.11834/jrs.20252301. [5] HARALICK R M, STERNBERG S R, and ZHUANG Xinhua. Image analysis using mathematical morphology[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1987, PAMI-9(4): 532–550. doi: 10.1109/TPAMI.1987.4767941. [6] HAO Congyu, LI Zhengzhou, ZHANG Yuting, et al. Infrared small target detection based on adaptive size estimation by multidirectional gradient filter[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007915. doi: 10.1109/TGRS.2024.3502421. [7] 张晶晶, 曹思华, 崔文楠, 等. 基于改进顶帽变换的红外弱小目标检测[J]. 电子与信息学报, 2024, 46(1): 267–276. doi: 10.11999/JEIT230892.ZHANG Jingjing, CAO Sihua, CUI Wennan, et al. Improved top-hat transform-based algorithm for infrared dim and small target detection[J]. Journal of Electronics & Information Technology, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562. doi: 10.11999/JEIT230892. [8] GAO Chenqiang, MENG Deyu, YANG Yi, et al. Infrared patch-image model for small target detection in a single image[J]. IEEE Transactions on Image Processing, 2013, 22(12): 4996–5009. doi: 10.1109/TIP.2013.2281420. [9] 李守昌, 吴滢跃. 复杂背景下双邻域局部权重对比度红外小目标检测算法[J]. 半导体光电, 2024, 45(5): 853–860. doi: 10.16818/j.issn1001-5868.2024040301.LI Shouchang and WU Yingyue. Dual-neighborhood local weighted contrast algorithm for infrared small target detection within complex backgrounds[J]. Semiconductor Optoelectronics, 2024, 45(5): 853–860. doi: 10.16818/j.issn1001-5868.2024040301. [10] 苟士淼, 刘兆瑜, 马鹏阁, 等. 基于IHBF的加权局部对比度红外小目标检测[J]. 电光与控制, 2025, 32(8): 86–91. doi: 10.3969/j.issn.1671-637X.2025.08.014.GOU Shimiao, LIU Zhaoyu, MA Pengge, et al. Weighted local contrast infrared small target detection based on IHBF[J]. Electronics Optics & Control, 2025, 32(8): 86–91. doi: 10.3969/j.issn.1671-637X.2025.08.014. [11] DAI Yimian and WU Yiquan. Reweighted infrared patch-tensor model with both nonlocal and local priors for single-frame small target detection[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2017, 10(8): 3752–3767. doi: 10.1109/JSTARS.2017.2700023. [12] 蹇渊, 黄自力, 王询. 基于随机化张量算法的红外弱小目标检测[J]. 激光技术, 2024, 48(1): 127–134. doi: 10.7510/jgjs.issn.1001-3806.2024.01.020.JIAN Yuan, HUANG Zili, and WANG Xun. Infrared small target detection based on randomized tensor algorithm[J]. Laser Technology, 2024, 48(1): 127–134. doi: 10.7510/jgjs.issn.1001-3806.2024.01.020. [13] SUN Yang, YANG Jungang, LONG Yunli, et al. Infrared small target detection via spatial-temporal total variation regularization and weighted tensor nuclear norm[J]. IEEE Access, 2019, 7: 56667–56682. doi: 10.1109/ACCESS.2019.2914281. [14] LIU Ting, YANG Jungang, LI Boyang, et al. Nonconvex tensor low-rank approximation for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60: 5614718. doi: 10.1109/TGRS.2021.3130310. [15] 胡亮, 杨德贵, 赵党军, 等. 改进非凸估计与非对称时空正则化的红外小目标检测方法[J]. 国防科技大学学报, 2024, 46(3): 180–194. doi: 10.11887/j.cn.202403018.HU Liang, YANG Degui, ZHAO Dangjun, et al. Infrared small target detection method based on improved non-convex estimation and asymmetric spatial-temporal regularization[J]. Journal of National University of Defense Technology, 2024, 46(3): 180–194. doi: 10.11887/j.cn.202403018. [16] LI Fenghong, RAO Peng, SUN Wen, et al. A new motion feature-enhanced multiframe spatial-temporal infrared target detection network[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5006819. doi: 10.1109/TGRS.2025.3603784. [17] XU Jielei, HAN Xinheng, WANG Jiacheng, et al. PCLC-Net: Parallel connected lateral chain networks for infrared small target detection[J]. Remote Sensing, 2025, 17(12): 2072. doi: 10.3390/rs17122072. [18] 龙畅, 张弦, 刘艳阳, 等. 天基红外小目标探测技术新进展、挑战与对策[J]. 光电工程, 2025, 52(12): 250300. doi: 10.12086/oee.2025.250300.LONG Chang, ZHANG Xian, LIU Yanyang, et al. Recent progress, challenges, and countermeasures in the detection of dim small targets using space-based infrared systems[J]. Opto-Electronic Engineering, 2025, 52(12): 250300. doi: 10.12086/oee.2025.250300. [19] LI Fenghong, RAO Peng, SUN Wen, et al. A low signal-to-noise ratio infrared small-target detection network[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 8643–8658. doi: 10.1109/JSTARS.2025.3550581. [20] LIU Ming, DU Haoyuan, ZHAO Yuejin, et al. Image small target detection based on deep learning with SNR controlled sample generation[J]. Current Trends in Computer Science and Mechanical Automation, 2017, 1(1): 211–220. doi: 10.1515/9783110584974-025. [21] HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition[C]. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, USA, 2016: 770–778. doi: 10.1109/CVPR.2016.90. [22] 盛卫东, 吴双林, 肖超, 等. 可微稀疏掩模引导的红外小目标快速检测网络[J]. 电子与信息学报, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989.SHENG Weidong, WU Shuanglin, XIAO Chao, et al. Differentiable sparse mask guided infrared small target fast detection network[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989. [23] ZHANG Mingjin, ZHANG Rui, YANG Yuxiang, et al. ISNet: Shape matters for infrared small target detection[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA, 2022: 867–876. doi: 10.1109/CVPR52688.2022.00095. [24] ZHANG Tianfang, LI Lei, CAO Siying, et al. Attention-guided pyramid context networks for detecting infrared small target under complex background[J]. IEEE Transactions on Aerospace and Electronic Systems, 2023, 59(4): 4250–4261. doi: 10.1109/TAES.2023.3238703. [25] LI Chengyu, ZHANG Yan, SHI Zhiguang, et al. Moderately dense adaptive feature fusion network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5616712. doi: 10.1109/TGRS.2024.3381006. [26] LIN Fanzhao, BAO Kexin, LI Yong, et al. Learning contrast-enhanced shape-biased representations for infrared small target detection[J]. IEEE Transactions on Image Processing, 2024, 33: 3047–3058. doi: 10.1109/TIP.2024.3391011. [27] ZHOU Zongwei, RAHMAN SIDDIQUEE M M, TAJBAKHSH N, et al. Unet++: A nested u-net architecture for medical image segmentation[C]. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support 4th International Workshop, Granada, Spain, 2018: 3–11. doi: 10.1007/978-3-030-00889-5_1 . [28] DAI Yimian, WU Yiquan, ZHOU Fei, et al. Asymmetric contextual modulation for infrared small target detection[C]. Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, USA, 2021: 949–958. doi: 10.1109/WACV48630.2021.00099 . [29] DAI Yimian, WU Yiquan, ZHOU Fei, et al. Attentional local contrast networks for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 59(11): 9813–9824. doi: 10.1109/TGRS.2020.3044958. [30] LI Boyang, XIAO Chao, WANG Longguang, et al. Dense nested attention network for infrared small target detection[J]. IEEE Transactions on Image Processing, 2023, 32: 1745–1758. doi: 10.1109/TIP.2022.3199107. [31] LIN T Y, DOLLÁR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, USA, 2017: 2117–2125. doi: 10.1109/CVPR.2017.106. [32] 高凯, 王晟宇, 付强, 等. 基于多层级交互式特征融合的三维目标检测算法[J]. 吉林大学学报: 理学版, 2026, 64(3): 591–602. doi: 10.13413/j.cnki.jdxblxb.2025021.GAO Kai, WANG Shengyu, FU Qiang, et al. Three-dimensional object detection algorithm based on multi-level interactive feature fusion[J]. Journal of Jilin University: Science Edition, 2026, 64(3): 591–602. doi: 10.13413/j.cnki.jdxblxb.2025021. [33] GUO Huinan, ZHANG Nengshuang, ZHANG Jing, et al. Location-guided dense nested attention network for infrared small target detection[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17: 18535–18548. doi: 10.1109/JSTARS.2024.3472041. [34] HU Jie, SHEN Li, and SUN Gang. Squeeze-and-excitation networks[C]. Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 7132–7141. doi: 10.1109/CVPR.2018.00745. [35] LI Xiang, WANG Wenhai, HU Xiaolin, et al. Selective kernel networks[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 510–519. doi: 10.1109/CVPR.2019.00060. [36] WANG Qilong, WU Banggu, ZHU Pengfei, et al. ECA-Net: Efficient channel attention for deep convolutional neural networks[C]. Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2020: 11531–11539. doi: 10.1109/CVPR42600.2020.01155. [37] WOO S H, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. 15th European Conference on Computer Vision-ECCV 2018, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1. [38] PARK J, WOO S, LEE J Y, et al. BAM: Bottleneck attention module[C]. Proceedings of the 2018 British Machine Vision Conference (BMVC), Newcastle, UK, 2018: 1–20. doi: 10.5244/C.32.1. (查阅网上资料,未找到本条文献doi信息,请确认). [39] LIU Yichao, SHAO Zongru, and HOFFMANN N. Global attention mechanism: Retain information to enhance channel-spatial interactions[EB/OL]. https://arxiv.org/abs/2112.05561, 2021. [40] ZHONG Yuhan, SHI Zhiguang, ZHANG Yan, et al. CSAN-UNET: Channel spatial attention nested UNet for infrared small target detection[J]. Remote Sensing, 2024, 16(11): 1894. doi: 10.3390/rs16111894. [41] ZUO Zhen, TONG Xiaozhong, WEI Junyu, et al. AFFPN: Attention fusion feature pyramid network for small infrared target detection[J]. Remote Sensing, 2022, 14(14): 3412. doi: 10.3390/rs14143412. [42] HOU Qibin, ZHOU Daquan, and FENG Jiashi. Coordinate attention for efficient mobile network design[C]. Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, USA, 2021: 13708–13717. doi: 10.1109/CVPR46437.2021.01350. [43] MISRA D, NALAMADA T, ARASANIPALAI A U, et al. Rotate to attend: Convolutional triplet attention module[C]. Proceedings of 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, USA, 2021: 3138–3147. doi: 10.1109/WACV48630.2021.00318. -
下载: