Advanced Search
Turn off MathJax
Article Contents
SHEN Xiaoru, CHANG Xia, WEI Wenjie. Dynamic Frequency Guidance and Semantic Purification Network for Infrared Dim and Small Target Detection[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260582
Citation: SHEN Xiaoru, CHANG Xia, WEI Wenjie. Dynamic Frequency Guidance and Semantic Purification Network for Infrared Dim and Small Target Detection[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260582

Dynamic Frequency Guidance and Semantic Purification Network for Infrared Dim and Small Target Detection

doi: 10.11999/JEIT260582 cstr: 32379.14.JEIT260582
Funds:  The Natural Science Foundation of Ningxia (2025AAC030002), The Construction Project of First-Class Disciplines in Ningxia Higher Education (NXYLXK2017B09), Graduate Innovation Program of North Minzu University for Nationality(CYX25121)
  • Received Date: 2026-05-11
  • Accepted Date: 2026-08-17
  • Rev Recd Date: 2026-08-14
  • Available Online: 2026-08-25
  •   Objective  Infrared dim and small target detection remains challenging because targets are small, have low contrast, and are easily obscured by complex background clutter. Although existing methods improve detection performance, several limitations remain, including feature loss during downsampling, insufficient use of physical priors, and contamination of semantic features by background noise. These limitations can result in missed detections and false alarms. To address these problems, a detection network combining dynamic frequency guidance and semantic purification is proposed. Deep semantic features are dynamically coupled with adaptive frequency-domain priors to strengthen target and edge representations, improve detection accuracy, reduce false alarms, and enhance robustness in complex scenes.  Methods  A framework consisting of a shared encoder and a dual-branch decoder is constructed for infrared dim and small target detection (Fig. 1). First, a Local Target Enhancement (LTA) module is used to preprocess the original infrared image, strengthen small-target boundaries, and suppress background noise. The Sobel operator is then applied to extract edge information and provide a structural prior for subsequent feature learning. During feature extraction, a Learnable Downsampling (LD) module replaces conventional max pooling. Stride-2 convolution adaptively compresses the feature maps while preserving target details and reducing information loss. Residual blocks and progressive multiscale feature fusion are also used to strengthen semantic representations at different levels. A Dynamic Frequency Guidance Module (DFGM) is then introduced to enhance high-frequency details and target-edge information. Deep semantic features are used to dynamically predict the cutoff radius of a high-pass filter. A content-adaptive high-pass filter mask is constructed in the frequency domain to suppress low-frequency components. The filtered features are subsequently transformed back to the spatial domain to generate a frequency guidance map. To suppress spurious background responses introduced during frequency-domain enhancement, a Semantic Purification Module (SPM) is further designed. The deepest semantic features are used to generate a spatial attention map for top-down unidirectional filtering of the frequency guidance features. Background interference is thereby suppressed without contaminating high-level semantic representations with low-level noise. Finally, the dual-branch decoder reconstructs the target region and edge structure, thereby improving segmentation consistency, target boundary localization, and target-shape recovery.  Results and Discussions  Experimental results on the SIRST-Aug and IRSTD-1k datasets show that the proposed method outperforms mainstream comparison methods. On SIRST-Aug (Table 1), the method achieves the best Iintersection over Union (IoU), normalized Intersection over Union (nIoU), area under the receiver operating characteristic curve (AUC), and probability of detection (Pd), reaching 76.66%, 73.20%, 94.45%, and 99.17%, respectively. The false-alarm rate (Fa) is maintained at 33.96 × 10–6. On IRSTD-1k (Table 1), the method achieves the best IoU, AUC, and Fa values of 67.17%, 89.52%, and 10.62 × 10–6, respectively, while obtaining an nIoU of 67.67% and a Pd of 92.59%. Qualitative comparisons further show that the proposed method provides more accurate target localization and more complete target-shape segmentation under complex background interference (Figs. 4 and 5). The model contains 2.94 M parameters, requires 14.68 G FLoating-point OPerations (FLOPs), and has an average inference time of 7.54 ms, providing a favorable balance between detection performance and computational cost (Table 2). Ablation experiments further confirm the contributions of the key modules (Table 3). Adding DFGM increases AUC from 93.17% to 94.60%. Introducing LD increases nIoU to 73.32%, and adding SPM further increases IoU to 76.66%. The loss-function ablation results also show that the combined mask and edge losses provide a better balance among segmentation accuracy, target boundary localization, and false-alarm suppression (Table 4). Overall, the proposed method provides strong detection performance and effective false-alarm suppression.  Conclusions  An infrared dim and small target detection network based on dynamic frequency guidance and semantic purification is proposed. Through the coordinated use of LD, DFGM, and SPM, target details are effectively preserved, high-frequency structural information is strengthened, and semantic representations are purified. Background interference is consequently suppressed, and target-boundary recovery is improved. Experiments on SIRST-Aug and IRSTD-1k demonstrate that the proposed method outperforms existing mainstream methods across multiple evaluation metrics and maintains strong robustness in complex scenes. Future work can incorporate richer physical priors and self-supervised learning strategies to further improve model generalization and robustness under challenging conditions.
  • loading
  • [1]
    龙畅, 张弦, 刘艳阳, 等. 天基红外小目标探测技术新进展、挑战与对策[J]. 光电工程, 2025, 52(12): 250300. doi: 10.12086/oee.2025.250300.

    LONG Chang, ZHANG Xian, LIU Yanyang, et al. Recent progress, challenges, and countermeasures in the detection of dim small targets using space-based infrared systems[J]. Opto-Electronic Engineering, 2025, 52(12): 250300. doi: 10.12086/oee.2025.250300.
    [2]
    杨德贵, 韩同欢, 胡亮, 等. 单帧红外弱小目标检测技术研究现状与展望[J]. 信号处理, 2024, 40(5): 887–906. doi: 10.16798/j.issn.1003-0530.2024.05.008.

    YANG Degui, HAN Tonghuan, HU Liang, et al. Research status and prospect of single frame infrared dim small target detection technology[J]. Journal of Signal Processing, 2024, 40(5): 887–906. doi: 10.16798/j.issn.1003-0530.2024.05.008.
    [3]
    刘杰, 刘书豪, 田明, 等. 复杂环境下无人机航拍小目标检测算法[J]. 电子与信息学报, 2026, 48(4): 1763–1773. doi: 10.11999/JEIT251126.

    LIU Jie, LIU Shuhao, TIAN Ming, et al. Small object detection algorithm for UAV aerial images in complex environments[J]. Journal of Electronics & Information Technology, 2026, 48(4): 1763–1773. doi: 10.11999/JEIT251126.
    [4]
    BAI Xiangzhi and ZHOU Fugen. Analysis of new top-hat transformation and the application for infrared dim small target detection[J]. Pattern Recognition, 2010, 43(6): 2145–2156. doi: 10.1016/j.patcog.2009.12.023.
    [5]
    HAN Jinhui, MORADI S, FARAMARZI I, et al. Infrared small target detection based on the weighted strengthened local contrast measure[J]. IEEE Geoscience and Remote Sensing Letters, 2021, 18(9): 1670–1674. doi: 10.1109/LGRS.2020.3004978.
    [6]
    ZHANG Landan and PENG Zhenming. Infrared small target detection based on partial sum of the tensor nuclear norm[J]. Remote Sensing, 2019, 11(4): 382. doi: 10.3390/rs11040382.
    [7]
    张晶晶, 曹思华, 崔文楠, 等. 基于改进顶帽变换的红外弱小目标检测[J]. 电子与信息学报, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562.

    ZHANG Jingjing, CAO Sihua, CUI Wennan, et al. Improved top-hat transform-based algorithm for infrared dim and small target detection[J]. Journal of Electronics & Information Technology, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562.
    [8]
    DAI Yimian, WU Yiquan, ZHOU Fei, et al. Asymmetric contextual modulation for infrared small target detection[C]. The 2021 IEEE Winter Conference on Applications of Computer Vision, Waikoloa, USA, 2021: 949–958. doi: 10.1109/WACV48630.2021.00099.
    [9]
    DAI Yimian, WU Yiquan, ZHOU Fei, et al. Attentional local contrast networks for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 59(11): 9813–9824. doi: 10.1109/TGRS.2020.3044958.
    [10]
    LI Boyang, XIAO Chao, WANG Longguang, et al. Dense nested attention network for infrared small target detection[J]. IEEE Transactions on Image Processing, 2023, 32: 1745–1758. doi: 10.1109/TIP.2022.3199107.
    [11]
    SUN Heng, BAI Junxiang, YANG Fan, et al. Receptive-field and direction induced attention network for infrared dim small target detection with a large-scale dataset IRDST[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 5000513. doi: 10.1109/TGRS.2023.3235150.
    [12]
    ZHANG Tianfang, LI Lei, CAO Siying, et al. Attention-guided pyramid context networks for detecting infrared small target under complex background[J]. IEEE Transactions on Aerospace and Electronic Systems, 2023, 59(4): 4250–4261. doi: 10.1109/TAES.2023.3238703.
    [13]
    YUAN Shuai, QIN Hanlin, YAN Xiang, et al. SCTransNet: Spatial-channel cross transformer network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5002615. doi: 10.1109/TGRS.2024.3383649.
    [14]
    LI Qiang, ZHANG Mingwei, YANG Zhigang, et al. Edge-guided perceptual network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5643510. doi: 10.1109/TGRS.2024.3471865.
    [15]
    HU Chen, HUANG Yian, LI Kexuan, et al. DATransNet: Dynamic attention transformer network for infrared small target detection[J]. IEEE Geoscience and Remote Sensing Letters, 2025, 22: 7001005. doi: 10.1109/LGRS.2025.3557021.
    [16]
    MA Qianwen, DENG Shangwei, LI Bincheng, et al. DWTFreqNet: Infrared small target detection via wavelet-driven frequency matching and saliency-difference optimization[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5007815. doi: 10.1109/TGRS.2025.3608725.
    [17]
    ZHANG Yingmei, BAO Wangtao, YANG Yong, et al. MPCNet: Multiscale perception and cross-attention feature fusion network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2026, 64: 5000915. doi: 10.1109/TGRS.2026.3653023.
    [18]
    汪佳旭, 杨俊, 许聪源. 基于YOLO的自适应多尺度红外目标检测网络[J]. 光电工程, 2026, 53(4): 250292. doi: 10.12086/oee.2026.250292.

    WANG Jiaxu, YANG Jun, and XU Congyuan. YOLO-based adaptive multi-scale infrared target detection network[J]. Opto-Electronic Engineering, 2026, 53(4): 250292. doi: 10.12086/oee.2026.250292.
    [19]
    盛卫东, 吴双林, 肖超, 等. 可微稀疏掩模引导的红外小目标快速检测网络[J]. 电子与信息学报, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989.

    SHENG Weidong, WU Shuanglin, XIAO Chao, et al. Differentiable sparse mask guided infrared small target fast detection network[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989.
    [20]
    SPRINGENBERG J T, DOSOVITSKIY A, BROX T, et al. Striving for simplicity: The all convolutional net[C]. International Conference on Learning Representations, San Diego, USA, 2015.
    [21]
    邹旻瑞, 李宇轩, 戴一冕, 等. UMM-Det: 面向异构多模态遥感影像的一体化目标检测框架[J]. 电子与信息学报, 2025, 47(12): 4704–4713. doi: 10.11999/JEIT250933.

    ZOU Minrui, LI Yuxuan, DAI Yimian, et al. UMM-Det: A unified object detection framework for heterogeneous multimodal remote sensing imagery[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4704–4713. doi: 10.11999/JEIT250933.
    [22]
    鞠默然, 罗海波, 刘广琦, 等. 采用空间注意力机制的红外弱小目标检测网络[J]. 光学 精密工程, 2021, 29(4): 843–853. doi: 10.37188/OPE.20212904.0843.

    JU Moran, LUO Haibo, LIU Guangqi, et al. Infrared dim and small target detection network based on spatial attention mechanism[J]. Optics and Precision Engineering, 2021, 29(4): 843–853. doi: 10.37188/OPE.20212904.0843.
    [23]
    HU Jie, SHEN Li, and SUN Gang. Squeeze-and-excitation networks[C]. The 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 7132–7141. doi: 10.1109/CVPR.2018.00745.
    [24]
    WOO S, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. The 15th European Conference on Computer Vision–ECCV 2018, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1.
    [25]
    ZHANG Zongjian, WU Qiang, WANG Yang, et al. High-quality image captioning with fine-grained and semantic-guided visual attention[J]. IEEE Transactions on Multimedia, 2019, 21(7): 1681–1693. doi: 10.1109/TMM.2018.2888822.
    [26]
    CHEN Zifa. LGI-DETR: Local-global interaction for UAV object detection[C]. The 21st International Conference on Intelligent Computing Technology and Applications, Ningbo, China, 2025: 42–53. doi: 10.1007/978-981-96-9901-8_4.
    [27]
    ZHANG Mingjin, ZHANG Rui, YANG Yuxiang, et al. ISNet: Shape matters for infrared small target detection[C]. The 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 867–876. doi: 10.1109/CVPR52688.2022.00095.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(5)  / Tables(4)

    Article Metrics

    Article views (95) PDF downloads(16) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return