Advanced Search
Turn off MathJax
Article Contents
LIU Minglong, JIANG Tingyao, LI Yulan. Frequency-Domain Decoupling and Spatial-Prior-Constrained Detection Method for Infrared Dim and Small Targets[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260847
Citation: LIU Minglong, JIANG Tingyao, LI Yulan. Frequency-Domain Decoupling and Spatial-Prior-Constrained Detection Method for Infrared Dim and Small Targets[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260847

Frequency-Domain Decoupling and Spatial-Prior-Constrained Detection Method for Infrared Dim and Small Targets

doi: 10.11999/JEIT260847 cstr: 32379.14.JEIT260847
Funds:  National Key R&D Program of China (2024YFD1702005), National Natural Science Foundation of China (Program of Joint Funds: U25A20403), Special Fund for Central Government Guiding Local Science and Technology Development of Hubei Province (2024BSB002)
  • Accepted Date: 2026-08-17
  • Rev Recd Date: 2026-08-17
  • Available Online: 2026-08-26
  •   Objective  Infrared dim and small target detection exploits thermal-radiation differences between targets and backgrounds for passive sensing in air-ground inspection, maritime surveillance, and wide-area warning. Under long-range imaging, low signal-to-noise ratios, and complex backgrounds, targets occupy few pixels and exhibit weak textures, blurred contours, and low contrast. Bounding-box detection is well suited to edge-deployed rapid detection and subsequent tracking initialization by directly predicting class confidence and location. Although the Real-Time Detection Transformer (RT-DETR) provides an efficient end-to-end framework, its application faces three limitations: repeated downsampling weakens local details and initial localization cues; intra-scale global interaction insufficiently distinguishes high-frequency target responses from low-frequency background context; and cross-scale fusion may propagate homogeneous clutter without explicit spatial constraints. Therefore, a Frequency-Domain Decoupling and Spatial-Prior-Constrained DETR (FSP-DETR) is proposed to coordinate shallow detail preservation, target-background decoupling, and localization constraints.  Methods  FSP-DETR is built on RT-DETR-R18 and follows backbone feature extraction, intra-scale interaction, cross-scale fusion, and query-based decoding (Fig.1). The three modules operate at successive stages to address the three identified limitations. First, a Subspace Progressive Attention Modulated Inverted Residual Block (SPA-MIRB) is embedded in the Residual Network 18 (ResNet18) backbone to preserve shallow details (Fig.2). Its inverted-residual branch combines pointwise and depthwise convolutions with a residual connection to retain edges, spots, and weak textures; its attention branch partitions channels into subspaces and progressively transfers interactions to strengthen weak targets. Second, a Frequency-Domain Asymmetrically Decoupled Attention-based Intra-scale Feature Interaction module (FD-AIFI) replaces the Attention-Based Intra-Scale Feature Interaction module (AIFI) (Fig.3). Unified global interaction is divided into high-pass local-attention and low-pass global-attention paths. The former uses local-window attention to preserve compact target responses and limit distant clutter. The latter retains query resolution but pools key and value features, compressing low-frequency context while preserving spatial queries and target-position indices. Their outputs are concatenated channel-wise and linearly projected. Third, a Spatial-Prior-Constrained CNN-based Cross-scale Feature-fusion Module (SPC-CCFM) retains the upsampling, concatenation, and convolution path while introducing a High-Order Spatial Representation Module (HSRM) and an Overall-Level Mapping Injection Mechanism (OMIM). HSRM aligns multilevel features and constructs spatial passthrough, high-order spatial-relation, and low-order fidelity branches (Fig.4). The branches preserve original spatial responses, model spatially separated yet response-similar clutter through hypergraph aggregation, and supplement local edges and weak textures. OMIM adapts and injects the prior into each fusion node to constrain cross-scale refinement.  Results and Discussions  Experiments are conducted on IRSTD-1k and NUAA-SIRST, with pixel-level masks converted into single-class minimum bounding rectangles. Component ablation shows that SPA-MIRB, FD-AIFI, and SPC-CCFM each improve mean Average Precision at an Intersection over Union threshold of 0.5 (mAP@0.5) on both datasets (Table 1). FSP-DETR achieves 89.03% and 98.75% mAP@0.5 on IRSTD-1k and NUAA-SIRST, exceeding RT-DETR-R18 by 3.39 and 3.08 percentage points. Parameters decrease from 19.87 million to 17.24 million and computational cost from 56.9 to 50.4 billion floating-point operations, by 13.2% and 11.4%. Although some variants yield higher precision or recall at one operating point, the complete model obtains the highest mAP@0.5 with lower complexity. Replacement experiments support the structural choices: SPA-MIRB and FD-AIFI achieve the highest mAP@0.5 among their counterparts, while SPC-CCFM exceeds both fusion alternatives by 0.48 and 0.26 percentage points, respectively, highlighting spatial constraints over complexity reduction (Table 2). SPA-MIRB reaches 96.75% mAP@0.5 with 15.28 million parameters and 46.8 billion floating-point operations, providing a balanced accuracy-complexity trade-off. A balanced path-allocation factor of 0.50 performs best, reaching 97.27% mAP@0.5 and 53.04% mAP@0.5:0.95 on NUAA-SIRST (Table 3). Attention maps show concentrated target-neighborhood responses in the high-pass path and broader, smoother background responses in the low-pass path (Fig.5). Under unified settings, FSP-DETR obtains the highest mAP@0.5 on both datasets, with 172 frames/s and an average per-image forward time of 5.81 ms (Table 4). It also surpasses deeper RT-DETR variants with fewer parameters, less computation, and shorter forward time, showing the advantage of task-oriented adaptation over backbone enlargement. Detection visualization shows better agreement between predicted boxes and target regions in cluttered scenes (Fig.6). Instance-level analysis shows missed-detection rates decrease from 15.70% to 13.22% on IRSTD-1k and from 10.71% to 5.36% on NUAA-SIRST, while average false positives per image decrease from 0.490 to 0.380 and from 0.209 to 0.070 (Table 5). Center-offset errors also decrease on both datasets. However, the scale error on NUAA-SIRST increases slightly from 10.63% to 11.05%, indicating that scale regression for extremely small targets remains challenging.  Conclusions  FSP-DETR coordinates shallow detail preservation, frequency-domain decoupled intra-scale interaction, and spatial-prior-constrained cross-scale fusion for end-to-end bounding-box detection. It improves accuracy, missed-detection control, false-alarm suppression, and center localization while reducing complexity and maintaining efficient inference. An accuracy-efficiency balance is achieved in complex infrared scenes. Future work will investigate multi-frame spatiotemporal modeling for thermal-crossover backgrounds and dense dim-target scenes.
  • loading
  • [1]
    ZHANG Nan, LIU Youmeng, LIU Hao, et al. DTNet: A specialized dual-tuning network for infrared vehicle detection in aerial images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5002815. doi: 10.1109/TGRS.2024.3386309.
    [2]
    LI Boyang, XIAO Chao, WANG Longguang, et al. Dense nested attention network for infrared small target detection[J]. IEEE Transactions on Image Processing, 2023, 32: 1745–1758. doi: 10.1109/TIP.2022.3199107.
    [3]
    LIN Fanzhao, BAO Kexin, LI Yong, et al. Learning contrast-enhanced shape-biased representations for infrared small target detection[J]. IEEE Transactions on Image Processing, 2024, 33: 3047–3058. doi: 10.1109/TIP.2024.3391011.
    [4]
    LI Fenghong, RAO Peng, SUN Wen, et al. A low signal-to-noise ratio infrared small-target detection network[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 8643–8658. doi: 10.1109/JSTARS.2025.3550581.
    [5]
    LI Qiang, ZHANG Mingwei, YANG Zhigang, et al. Edge-guided perceptual network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5643510. doi: 10.1109/TGRS.2024.3471865.
    [6]
    YANG Huoren, MU Tingkui, DONG Ziyue, et al. PBT: Progressive background-aware transformer for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5004513. doi: 10.1109/TGRS.2024.3415080.
    [7]
    ZHANG Shizhou, WANG Zhang, XING Yinghui, et al. SCAFNet: Semantic-guided cascade adaptive fusion network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007712. doi: 10.1109/TGRS.2024.3492256.
    [8]
    张晶晶, 曹思华, 崔文楠, 等. 基于改进顶帽变换的红外弱小目标检测[J]. 电子与信息学报, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562.

    ZHANG Jingjing, CAO Sihua, CUI Wennan, et al. Improved top-hat transform-based algorithm for infrared dim and small target detection[J]. Journal of Electronics & Information Technology, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562.
    [9]
    薛驰, 陈小梅, 李海彤. 融合背景估计与相对局部对比度的天基短波红外弱小目标检测[J]. 光学 精密工程, 2026, 34(3): 450–465. doi: 10.37188/OPE.20263403.0450.

    XUE Chi, CHEN Xiaomei, and LI Haitong. SWIR weak targets detection on space-based platform integrating background estimation and relative local contrast[J]. Optics and Precision Engineering, 2026, 34(3): 450–465. doi: 10.37188/OPE.20263403.0450.
    [10]
    YANG Bo, ZHANG Xinyu, ZHANG Jian, et al. EFLNet: Enhancing feature learning network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5906511. doi: 10.1109/TGRS.2024.3365677.
    [11]
    CHEN Tianxiang, TAN Zhentao, GONG Tao, et al. Feature preservation and shape cues assist infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5006412. doi: 10.1109/TGRS.2024.3461795.
    [12]
    盛卫东, 吴双林, 肖超, 等. 可微稀疏掩模引导的红外小目标快速检测网络[J]. 电子与信息学报, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989.

    SHENG Weidong, WU Shuanglin, XIAO Chao, et al. Differentiable sparse mask guided infrared small target fast detection network[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989.
    [13]
    吕鹏远, 兰金江, 曾学仁, 等. 基于特征增强与融合的红外目标检测算法[J]. 红外技术, 2024, 46(7): 782–790.

    LYU Pengyuan, LAN Jinjiang, ZENG Xueren, et al. Infrared object detection algorithm based on feature enhancement and fusion[J]. Infrared Technology, 2024, 46(7): 782–790.
    [14]
    YUAN Shuai, QIN Hanlin, YAN Xiang, et al. SCTransNet: Spatial-channel cross transformer network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5002615. doi: 10.1109/TGRS.2024.3383649.
    [15]
    CHEN Tianxiang, YE Zi, TAN Zhentao, et al. MiM-ISTD: Mamba-in-mamba for efficient infrared small-target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007613. doi: 10.1109/TGRS.2024.3485721.
    [16]
    MA Tianfeng, GUO Guanqin, LI Zhimeng, et al. Infrared small target detection method based on high-low-frequency semantic reconstruction[J]. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 6012505. doi: 10.1109/LGRS.2024.3428626.
    [17]
    XU Mingzhu, YU Chenglong, LI Zexuan, et al. HDNet: A hybrid domain network with multiscale high-frequency information enhancement for infrared small-target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5004115. doi: 10.1109/TGRS.2025.3574962.
    [18]
    刘皓皎, 刘力双, 张明淳. 基于YOLOv5改进的红外目标检测算法[J]. 激光技术, 2024, 48(4): 534–541. doi: 10.7510/jgjs.issn.1001-3806.2024.04.011.

    LIU Haojiao, LIU Lishuang, and ZHANG Mingchun. An improved infrared object detection algorithm based on YOLOv5[J]. Laser Technology, 2024, 48(4): 534–541. doi: 10.7510/jgjs.issn.1001-3806.2024.04.011.
    [19]
    汪佳旭, 杨俊, 许聪源. 基于YOLO的自适应多尺度红外目标检测网络[J]. 光电工程, 2026, 53(4): 250292. doi: 10.12086/oee.2026.250292.

    WANG Jiaxu, YANG Jun, and XU Congyuan. YOLO-based adaptive multi-scale infrared target detection network[J]. Opto-Electronic Engineering, 2026, 53(4): 250292. doi: 10.12086/oee.2026.250292.
    [20]
    王坤, 丁麒龙. 利用自适应融合和混合锚检测器的遥感图像小目标检测算法[J]. 电子与信息学报, 2024, 46(7): 2942–2951. doi: 10.11999/JEIT230966.

    WANG Kun and DING Qilong. Remote sensing images small object detection algorithm with adaptive fusion and hybrid anchor detector[J]. Journal of Electronics & Information Technology, 2024, 46(7): 2942–2951. doi: 10.11999/JEIT230966.
    [21]
    刘杰, 刘书豪, 田明, 等. 复杂环境下无人机航拍小目标检测算法[J]. 电子与信息学报, 2026, 48(4): 1763–1773. doi: 10.11999/JEIT251126.

    LIU Jie, LIU Shuhao, TIAN Ming, et al. Small object detection algorithm for UAV aerial images in complex environments[J]. Journal of Electronics & Information Technology, 2026, 48(4): 1763–1773. doi: 10.11999/JEIT251126.
    [22]
    XIAO Bin and HU Yue. Bounding box regression network for infrared small target detection with adaptive receptive field and cross-scale fusion[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5003514. doi: 10.1109/TGRS.2025.3564958.
    [23]
    张上, 崔玉杰, 龚国强, 等. LIST-DETR: 复杂背景下红外小目标检测的高效轻量级检测网络[J]. 红外与激光工程, 2025, 54(10): 20250271. doi: 10.3788/IRLA20250271.

    ZHANG Shang, CUI Yujie, GONG Guoqiang, et al. LIST-DETR: An efficient and lightweight detection network for infrared small target detection in complex backgrounds[J]. Infrared and Laser Engineering, 2025, 54(10): 20250271. doi: 10.3788/IRLA20250271.
    [24]
    ZHAO Yian, LV Wenyu, XU Shangliang, et al. DETRs beat YOLOs on real-time object detection[C]. Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 16965–16974. doi: 10.1109/CVPR52733.2024.01605.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(6)  / Tables(5)

    Article Metrics

    Article views (100) PDF downloads(10) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return