Dynamic Frequency Guided and Semantic Purification Network for Infrared Dim and Small Target Detection
-
摘要: 红外弱小目标检测因目标尺寸微小、信噪比低及背景杂波复杂而极具挑战性。因此,该文提出一种基于动态频率引导与语义净化的红外弱小目标检测方法。针对弱小目标在深度神经网络中易被淹没的问题,通过可学习的步进卷积替代传统的最大池化层,有效保留编码过程中的目标细节。其次,设计了动态频率引导模块,使用小型网络动态预测高通滤波截止半径,实现对不同场景下目标的自适应频率引导。为进一步抑制背景杂波,提出了单向语义净化模块,利用高层语义特征生成空间注意力权重,对频率引导图进行过滤和强化,从而突出目标细节并增强模型的感知能力。在SIRST-Aug和IRSTD-1k两个公开数据集上的实验结果表明,所提方法性能优于现有的主流方法,特别是在SIRST-Aug数据集上的检测率Pd达到了99.17 %,有效克服了红外弱小目标检测中误检与漏检的问题。Abstract:
Objective Infrared dim and small target detection faces many challenges, such as small target size, low contrast and complex background interference. Although the existing methods improve the detection performance, there are still some problems such as the loss of down-sampling features, the insufficient utilization of physical priors, and the pollution of semantic features by background noise, which lead to missed detection and false positives. To this end, this paper proposes a detection network combining dynamic frequency guidance and semantic purification, which combines deep semantic features with adaptive physical frequency priors to enhance target and edge representations, achieve higher detection accuracy, reduce false alarms, and improve the robustness of the model in complex scenes. Methods This paper presents a framework for infrared dim and small target detection based on shared encoder and dual decoders ( Fig.1 ). Firstly, the LTA module is used to preprocess the original infrared image to enhance the boundaries of small targets, suppress background noise, and improve the separability of the target. Subsequently, the Sobel operator is introduced to extract edge information, providing a structural prior for subsequent feature learning. In the feature extraction stage, the encoder builds the LD module, replacing the traditional maximum pooling operation and achieving adaptive feature compression through convolution, thereby reducing the spatial resolution while retaining the target's detailed information. At the same time, the network enhances its modeling ability for different semantic information through residual blocks and progressive multi-scale feature fusion mechanisms. To further improve the response to the target area, DFGM module is designed to enhance high-frequency details and target edge information. Additionally, SPM module is proposed, using the spatial attention mechanism to suppress background interference and optimize the purity of feature representation. Finally, a dual-branch collaborative decoder is adopted to restore and reconstruct the target area and edge structure, thereby improving the consistency of the segmentation result and enhancing the accuracy of object boundary positioning.Results and Discussions Experimental results on SIRST-AUG and IRSTD-1K datasets show that the proposed method is superior to the existing mainstream methods. On the SIRST-AUG dataset ( Table 1 ), the method in this paper achieved the best IoU, nIoU, AUC, and Pd, which were 76.66, 73.20, 94.45, and 99.17 respectively, and Fa was controlled at a low level of 33.96. On the IRSTD-1K dataset (Table 1 ), the proposed method achieved the best IoU (67.17), AUC (89.52), and Fa(10.62), while maintaining good nIoU and Pd Quantitative analysis further proves that this method can achieve more accurate target localization and more complete target shape segmentation in complex background scenarios(Fig.4 ,Fig.5 ). The proposed algorithm has moderate parameter quantity and an average inference time of 7.54 ms, effectively balancing the model's performance and computational cost(Table 2 ). Additionally, ablation experiments further validate the effectiveness of key modules (Table 3 ): DFGM module increased AUC from 93.17 to 94.60, the learnable downsampling strategy increased nIoU to 73.32, and SPM module increased IoU to 76.66. At the same time, the ablation experiments of the loss function demonstrate the effectiveness and robustness of the proposed algorithm architecture(Table 4 ). In conclusion, this method has strong detection performance and false alarm suppression ability.Conclusions This paper proposes an infrared dim and small target detection network based on dynamic frequency guidance and semantic purification. Through the collaborative action of local target enhancement, learnable downsampling, DFGM and SPM modules, it effectively retains the fine texture of the target, enhances the high-frequency information and semantic feature expression, while suppressing background interference and significantly improving the recovery quality of the target boundary. Experimental results on the SIRST-AUG and IRSTD-1K datasets show that this method outperforms existing mainstream methods in multiple key evaluation indicators and demonstrates strong robustness in complex scenarios. Future research can further combine more abundant physical prior knowledge and self-supervised learning strategies to enhance the generalization ability and robustness of the model in complex scenarios. -
表 1 不同算法在SIRST-AUG和IRSTD-1K数据集上的定量比较
Models SIRST-AUG IRSTD-1K IoU(%) nIoU(%) AUC(%) Pd(%) Fa($ {10}^{-6} $) IoU(%) nIoU(%) AUC(%) Pd(%) Fa($ {10}^{-6} $) TopHat[4] 17.02 22.39 59.11 82.94 133.77 6.08 19.52 60.87 75.42 716.40 WSLCM[5] 5.35 10.95 52.70 67.68 18.70 10.60 19.06 55.53 62.96 11.62 PSTNN[6] 20.77 28.31 60.78 63.27 71.48 17.35 21.36 60.72 65.32 65.21 ACM[8] 65.06 65.44 87.61 93.67 74.92 61.68 58.23 87.94 90.23 19.81 ALCNet[9] 66.30 66.99 86.56 92.15 42.39 60.80 61.14 85.09 85.85 26.84 DNANet[10] 72.78 70.54 89.79 96.83 35.87 65.74 64.85 80.54 89.22 27.27 RDIAN[11] 70.51 69.64 92.73 97.24 87.41 63.69 65.65 86.19 90.90 18.98 AGPCNet[12] 72.10 69.98 92.81 98.07 25.84 64.68 61.08 86.40 87.87 14.63 SCTransNet[13] 68.24 67.67 92.81 95.87 46.64 66.13 66.66 85.90 92.59 11.18 EGPNet[14] 76.37 72.39 93.17 99.03 28.10 66.97 67.98 89.50 93.93 17.57 DATransNet[15] 68.41 67.11 93.96 93.26 51.68 66.81 66.70 88.03 90.23 13.15 本文方法 76.66 73.20 94.45 99.17 33.96 67.17 67.67 89.52 92.59 10.62 表 2 不同算法的参数量、运算量和耗时对比
表 3 模块消融实验定量对比结果
Module SIRST-AUG DFGM LD SPM IoU(%) nIoU(%) AUC(%) × × × 76.37 72.39 93.17 √ × × 75.97 72.72 94.60 √ √ × 76.22 73.32 93.69 √ √ √ 76.66 73.20 94.45 表 4 损失函数消融实验定量对比结果
Lmask Ledge SIRST-AUG SoftIoU BCE SoftIoU IoU(%) nIoU(%) AUC(%) Fa(10–6) √ × × 74.47 71.32 94.44 71.56 √ × √ 73.65 71.76 91.90 27.24 √ √ × 76.78 73.18 94.63 55.07 √ √ √ 76.66 73.20 94.45 33.96 -
[1] 龙畅, 张弦, 刘艳阳, 等. 天基红外小目标探测技术新进展、挑战与对策[J]. 光电工程, 2025, 52(12): 250300. doi: 10.12086/oee.2025.250300.LONG Chang, ZHANG Xian, LIU Yanyang, et al. Recent progress, challenges, and countermeasures in the detection of dim small targets using space-based infrared systems[J]. Opto-Electronic Engineering, 2025, 52(12): 250300. doi: 10.12086/oee.2025.250300. [2] 杨德贵, 韩同欢, 胡亮, 等. 单帧红外弱小目标检测技术研究现状与展望[J]. 信号处理, 2024, 40(5): 887–906. doi: 10.16798/j.issn.1003-0530.2024.05.008.YANG Degui, HAN Tonghuan, HU Liang, et al. Research status and prospect of single frame infrared dim small target detection technology[J]. Journal of Signal Processing, 2024, 40(5): 887–906. doi: 10.16798/j.issn.1003-0530.2024.05.008. [3] 刘杰, 刘书豪, 田明, 等. 复杂环境下无人机航拍小目标检测算法[J]. 电子与信息学报, 2026, 48(4): 1763–1773. doi: 10.11999/JEIT251126.LIU Jie, LIU Shuhao, TIAN Ming, et al. Small object detection algorithm for UAV aerial images in complex environments[J]. Journal of Electronics & Information Technology, 2026, 48(4): 1763–1773. doi: 10.11999/JEIT251126. [4] BAI Xiangzhi and ZHOU Fugen. Analysis of new top-hat transformation and the application for infrared dim small target detection[J]. Pattern Recognition, 2010, 43(6): 2145–2156. doi: 10.1016/j.patcog.2009.12.023. [5] HAN Jinhui, MORADI S, FARAMARZI I, et al. Infrared small target detection based on the weighted strengthened local contrast measure[J]. IEEE Geoscience and Remote Sensing Letters, 2021, 18(9): 1670–1674. doi: 10.1109/LGRS.2020.3004978. [6] ZHANG Landan and PENG Zhenming. Infrared small target detection based on partial sum of the tensor nuclear norm[J]. Remote Sensing, 2019, 11(4): 382. doi: 10.3390/rs11040382. [7] 张晶晶, 曹思华, 崔文楠, 等. 基于改进顶帽变换的红外弱小目标检测[J]. 电子与信息学报, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562.ZHANG Jingjing, CAO Sihua, CUI Wennan, et al. Improved top-hat transform-based algorithm for infrared dim and small target detection[J]. Journal of Electronics & Information Technology, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562. [8] DAI Yimian, WU Yiquan, ZHOU Fei, et al. Asymmetric contextual modulation for infrared small target detection[C]. Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision, Waikoloa, USA, 2021: 949–958. doi: 10.1109/WACV48630.2021.00099. [9] DAI Yimian, WU Yiquan, ZHOU Fei, et al. Attentional local contrast networks for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 59(11): 9813–9824. doi: 10.1109/TGRS.2020.3044958. [10] LI Boyang, XIAO Chao, WANG Longguang, et al. Dense nested attention network for infrared small target detection[J]. IEEE Transactions on Image Processing, 2023, 32: 1745–1758. doi: 10.1109/TIP.2022.3199107. [11] SUN Heng, BAI Junxiang, YANG Fan, et al. Receptive-field and direction induced attention network for infrared dim small target detection with a large-scale dataset IRDST[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 5000513. doi: 10.1109/TGRS.2023.3235150. [12] ZHANG Tianfang, LI Lei, CAO Siying, et al. Attention-guided pyramid context networks for detecting infrared small target under complex background[J]. IEEE Transactions on Aerospace and Electronic Systems, 2023, 59(4): 4250–4261. doi: 10.1109/TAES.2023.3238703. [13] YUAN Shuai, QIN Hanlin, YAN Xiang, et al. SCTransNet: Spatial-channel cross transformer network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5002615. doi: 10.1109/TGRS.2024.3383649. [14] LI Qiang, ZHANG Mingwei, YANG Zhigang, et al. Edge-guided perceptual network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5643510. doi: 10.1109/TGRS.2024.3471865. [15] HU Chen, HUANG Yian, LI Kexuan, et al. DATransNet: Dynamic attention transformer network for infrared small target detection[J]. IEEE Geoscience and Remote Sensing Letters, 2025, 22: 7001005. doi: 10.1109/LGRS.2025.3557021. [16] MA Qianwen, DENG Shangwei, LI Bincheng, et al. DWTFreqNet: Infrared small target detection via wavelet-driven frequency matching and saliency-difference optimization[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5007815. doi: 10.1109/TGRS.2025.3608725. [17] ZHANG Yingmei, BAO Wangtao, YANG Yong, et al. MPCNet: Multiscale perception and cross-attention feature fusion network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2026, 64: 5000915. doi: 10.1109/TGRS.2026.3653023. [18] 汪佳旭, 杨俊, 许聪源. 基于YOLO的自适应多尺度红外目标检测网络[J]. 光电工程, 2026, 53(4): 250292. doi: 10.12086/oee.2026.250292.WANG Jiaxu, YANG Jun, and XU Congyuan. YOLO-based adaptive multi-scale infrared target detection network[J]. Opto-Electronic Engineering, 2026, 53(4): 250292. doi: 10.12086/oee.2026.250292. [19] 盛卫东, 吴双林, 肖超, 等. 可微稀疏掩模引导的红外小目标快速检测网络[J]. 电子与信息学报, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989.SHENG Weidong, WU Shuanglin, XIAO Chao, et al. Differentiable sparse mask guided infrared small target fast detection network[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989. [20] SPRINGENBERG J T, DOSOVITSKIY A, BROX T, et al. Striving for simplicity: The all convolutional net[C]. International Conference on Learning Representations, San Diego, USA, 2015. [21] 邹旻瑞, 李宇轩, 戴一冕, 等. UMM-Det: 面向异构多模态遥感影像的一体化目标检测框架[J]. 电子与信息学报, 2025, 47(12): 4704–4713. doi: 10.11999/JEIT250933.ZOU Minrui, LI Yuxuan, DAI Yimian, et al. UMM-Det: A unified object detection framework for heterogeneous multimodal remote sensing imagery[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4704–4713. doi: 10.11999/JEIT250933. [22] 鞠默然, 罗海波, 刘广琦, 等. 采用空间注意力机制的红外弱小目标检测网络[J]. 光学 精密工程, 2021, 29(4): 843–853. doi: 10.37188/OPE.20212904.0843.JU Moran, LUO Haibo, LIU Guangqi, et al. Infrared dim and small target detection network based on spatial attention mechanism[J]. Optics and Precision Engineering, 2021, 29(4): 843–853. doi: 10.37188/OPE.20212904.0843. [23] HU Jie, SHEN Li, and SUN Gang. Squeeze-and-excitation networks[C]. Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 7132–7141. doi: 10.1109/CVPR.2018.00745. [24] WOO S, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. Proceedings of the 15th European Conference on Computer Vision–ECCV 2018, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1. [25] ZHANG Zongjian, WU Qiang, WANG Yang, et al. High-quality image captioning with fine-grained and semantic-guided visual attention[J]. IEEE Transactions on Multimedia, 2019, 21(7): 1681–1693. doi: 10.1109/TMM.2018.2888822. [26] CHEN Zifa. LGI-DETR: Local-global interaction for UAV object detection[C]. Proceedings of the 21st International Conference on Intelligent Computing Technology and Applications, Ningbo, China, 2025: 42–53. doi: 10.1007/978-981-96-9901-8_4. [27] ZHANG Mingjin, ZHANG Rui, YANG Yuxiang, et al. ISNet: Shape matters for infrared small target detection[C]. Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 867–876. doi: 10.1109/CVPR52688.2022.00095. -
下载: