Frequency-Domain Decoupling and Spatial-Prior-Constrained Detection Method for Infrared Dim and Small Targets
-
摘要: 针对红外弱小目标检测中目标尺度小、纹理弱导致的细节表征衰减、低频背景渗透,以及实时检测变换器(RT-DETR)在跨尺度融合中易传播背景冗余、空间约束不足的问题,提出一种频域解耦与空间先验约束的端到端检测模型。该模型以子空间渐进注意力补偿浅层局部细节,以频域非对称解耦交互分离目标细节与背景低频成分,并通过高阶空间关系建模与空间先验注入增强跨尺度融合,约束同质杂波关联。在IRSTD-1k和NUAA-SIRST数据集上,所提模型较RT-DETR基准的mAP@0.5分别提升3.39和3.08个百分点,参数量和计算量分别降低13.24%和11.42%。结果表明,该模型能在降低开销的同时提升检测精度与定位鲁棒性。Abstract:
Objective Infrared dim and small target detection exploits thermal-radiation differences between targets and backgrounds for passive sensing in air-ground inspection, maritime surveillance, and wide-area warning. Under long-range imaging, low signal-to-noise ratios, and complex backgrounds, targets occupy few pixels and exhibit weak textures, blurred contours, and low contrast. Bounding-box detection is well suited to edge-deployed rapid detection and subsequent tracking initialization by directly predicting class confidence and location. Although the Real-Time Detection Transformer (RT-DETR) provides an efficient end-to-end framework, its application faces three limitations: repeated downsampling weakens local details and initial localization cues; intra-scale global interaction insufficiently distinguishes high-frequency target responses from low-frequency background context; and cross-scale fusion may propagate homogeneous clutter without explicit spatial constraints. Therefore, a Frequency-Domain Decoupling and Spatial-Prior-Constrained DETR (FSP-DETR) is proposed to coordinate shallow detail preservation, target-background decoupling, and localization constraints. Methods FSP-DETR is built on RT-DETR-R18 and follows backbone feature extraction, intra-scale interaction, cross-scale fusion, and query-based decoding ( Fig.1 ). The three modules operate at successive stages to address the three identified limitations. First, a Subspace Progressive Attention Modulated Inverted Residual Block (SPA-MIRB) is embedded in the Residual Network 18 (ResNet18) backbone to preserve shallow details (Fig.2 ). Its inverted-residual branch combines pointwise and depthwise convolutions with a residual connection to retain edges, spots, and weak textures; its attention branch partitions channels into subspaces and progressively transfers interactions to strengthen weak targets. Second, a Frequency-Domain Asymmetrically Decoupled Attention-based Intra-scale Feature Interaction module (FD-AIFI) replaces the Attention-Based Intra-Scale Feature Interaction module (AIFI) (Fig.3 ). Unified global interaction is divided into high-pass local-attention and low-pass global-attention paths. The former uses local-window attention to preserve compact target responses and limit distant clutter. The latter retains query resolution but pools key and value features, compressing low-frequency context while preserving spatial queries and target-position indices. Their outputs are concatenated channel-wise and linearly projected. Third, a Spatial-Prior-Constrained CNN-based Cross-scale Feature-fusion Module (SPC-CCFM) retains the upsampling, concatenation, and convolution path while introducing a High-Order Spatial Representation Module (HSRM) and an Overall-Level Mapping Injection Mechanism (OMIM). HSRM aligns multilevel features and constructs spatial passthrough, high-order spatial-relation, and low-order fidelity branches (Fig.4 ). The branches preserve original spatial responses, model spatially separated yet response-similar clutter through hypergraph aggregation, and supplement local edges and weak textures. OMIM adapts and injects the prior into each fusion node to constrain cross-scale refinement.Results and Discussions Experiments are conducted on IRSTD-1k and NUAA-SIRST, with pixel-level masks converted into single-class minimum bounding rectangles. Component ablation shows that SPA-MIRB, FD-AIFI, and SPC-CCFM each improve mean Average Precision at an Intersection over Union threshold of 0.5 (mAP@0.5) on both datasets ( Table 1 ). FSP-DETR achieves 89.03% and 98.75% mAP@0.5 on IRSTD-1k and NUAA-SIRST, exceeding RT-DETR-R18 by 3.39 and 3.08 percentage points. Parameters decrease from 19.87 million to 17.24 million and computational cost from 56.9 to 50.4 billion floating-point operations, by 13.2% and 11.4%. Although some variants yield higher precision or recall at one operating point, the complete model obtains the highest mAP@0.5 with lower complexity. Replacement experiments support the structural choices: SPA-MIRB and FD-AIFI achieve the highest mAP@0.5 among their counterparts, while SPC-CCFM exceeds both fusion alternatives by 0.48 and 0.26 percentage points, respectively, highlighting spatial constraints over complexity reduction (Table 2 ). SPA-MIRB reaches 96.75% mAP@0.5 with 15.28 million parameters and 46.8 billion floating-point operations, providing a balanced accuracy-complexity trade-off. A balanced path-allocation factor of 0.50 performs best, reaching 97.27% mAP@0.5 and 53.04% mAP@0.5:0.95 on NUAA-SIRST (Table 3 ). Attention maps show concentrated target-neighborhood responses in the high-pass path and broader, smoother background responses in the low-pass path (Fig.5 ). Under unified settings, FSP-DETR obtains the highest mAP@0.5 on both datasets, with 172 frames/s and an average per-image forward time of 5.81 ms (Table 4 ). It also surpasses deeper RT-DETR variants with fewer parameters, less computation, and shorter forward time, showing the advantage of task-oriented adaptation over backbone enlargement. Detection visualization shows better agreement between predicted boxes and target regions in cluttered scenes (Fig.6 ). Instance-level analysis shows missed-detection rates decrease from 15.70% to 13.22% on IRSTD-1k and from 10.71% to 5.36% on NUAA-SIRST, while average false positives per image decrease from 0.490 to 0.380 and from 0.209 to 0.070 (Table 5 ). Center-offset errors also decrease on both datasets. However, the scale error on NUAA-SIRST increases slightly from 10.63% to 11.05%, indicating that scale regression for extremely small targets remains challenging.Conclusions FSP-DETR coordinates shallow detail preservation, frequency-domain decoupled intra-scale interaction, and spatial-prior-constrained cross-scale fusion for end-to-end bounding-box detection. It improves accuracy, missed-detection control, false-alarm suppression, and center localization while reducing complexity and maintaining efficient inference. An accuracy-efficiency balance is achieved in complex infrared scenes. Future work will investigate multi-frame spatiotemporal modeling for thermal-crossover backgrounds and dense dim-target scenes. -
表 1 IRSTD-1k与NUAA-SIRST数据集消融实验
数据集 模型 A B C P/% R/% mAP@0.5/% F1/% 参数量/M 计算量/G IRSTD-1k RT-DETR — — — 87.02 84.38 85.64 85.68 19.87 56.9 A √ — — 91.51 85.68 88.35 88.50 15.28 46.8 B — √ — 92.26 82.78 87.97 87.26 19.83 57.1 C — — √ 87.99 83.75 86.98 86.35 21.86 60.4 D √ √ — 89.53 84.77 88.85 87.09 15.24 46.9 E √ — √ 87.77 84.78 87.22 86.42 17.27 50.3 F — √ √ 92.57 82.57 86.35 87.29 21.82 60.5 FSP-DETR √ √ √ 88.82 86.75 89.03 87.78 17.24 50.4 NUAA-SIRST RT-DETR — — — 93.22 89.29 95.67 91.21 19.87 56.9 A √ — — 98.51 87.50 96.75 92.68 15.28 46.8 B — √ — 96.29 92.66 97.27 94.44 19.83 57.1 C — — √ 94.61 89.29 96.17 91.87 21.86 60.4 D √ √ — 93.02 95.19 97.86 94.09 15.24 46.9 E √ — √ 92.71 91.07 95.11 91.88 17.27 50.3 F — √ √ 94.44 92.86 98.24 93.64 21.82 60.5 FSP-DETR √ √ √ 96.33 93.72 98.75 95.01 17.24 50.4 表 2 不同模块结构的对比实验结果
位置 模块 P/% R/% mAP@0.5/% 参数量/M 计算量/G 基准 RT-DETR 93.22 89.29 95.67 19.87 56.9 浅层细节保持 FasterBlock 97.40 85.71 95.84 16.79 49.5 浅层细节保持 Conv3XC 97.42 87.50 94.60 19.87 56.9 浅层细节保持 SPA-MIRB 98.51 87.50 96.75 15.28 46.8 尺度内交互 Pola 98.06 90.37 97.21 20.04 57.2 尺度内交互 AIFI-DAttention 94.57 93.24 96.30 19.88 57.2 尺度内交互 FD-AIFI 96.29 92.66 97.27 19.83 57.1 跨尺度融合 BiFPN 93.51 88.58 95.69 20.43 58.3 跨尺度融合 SDFM 94.08 88.94 95.91 21.24 59.4 跨尺度融合 SPC-CCFM 94.61 89.29 96.17 21.86 60.4 表 3 算力分配因子$ a $消融实验结果
数据集 模型 $ a $ P/% R/% F1/% mAP@0.5/% mAP@0.5:0.95/% NUAA-SIRST RT-DETR+FD-AIFI 0.00 95.23 89.29 92.16 93.65 50.34 0.25 91.18 92.31 91.74 95.21 48.04 0.50 96.29 92.66 94.44 97.27 53.04 0.75 91.61 91.07 91.34 96.17 49.64 1.00 94.36 89.55 91.89 96.27 50.92 表 4 不同类型检测方法性能与效率对比实验
模型 参数量
/M计算量
/G帧率
(f/s)前向耗时
/msIRSTD-1k NUAA-SIRST P/% R/% mAP@0.5/% P/% R/% mAP@0.5/% SCTransNet* 20.20 58.9 — — 87.20 81.60 85.10 96.10 91.70 96.90 YOLOv5m 25.04 64.0 185 5.41 84.27 83.44 85.03 96.07 92.86 95.44 YOLOv8m 25.84 78.7 165 6.06 82.19 82.78 83.28 94.28 93.64 94.72 YOLOv9m 20.01 76.5 170 5.88 85.40 81.34 85.08 96.15 89.13 92.68 YOLOv10m 15.31 58.9 210 4.76 83.00 80.82 84.09 96.12 88.48 93.49 YOLO11m 20.03 67.6 180 5.56 82.69 79.47 82.30 95.83 92.86 93.96 YOLO12m 20.10 67.1 182 5.49 84.58 83.52 86.97 95.32 89.29 91.56 YOLO26m 20.35 67.8 178 5.62 88.64 76.38 85.72 87.92 82.14 88.82 RT-DETR-R18 19.87 56.9 160 6.25 87.02 84.38 85.64 93.22 89.29 95.67 RT-DETR-R34 31.11 88.8 125 8.00 90.91 83.44 87.70 93.31 91.07 95.11 RT-DETR-R50 41.96 129.6 90 11.11 86.11 84.11 85.34 94.48 91.72 96.17 RT-DETR-L 31.99 103.5 115 8.70 88.15 83.79 84.26 98.11 92.79 97.70 LIST-DETR* 14.60 41.4 — — 89.60 83.00 85.80 97.80 92.10 97.20 FSP-DETR 17.24 50.4 172 5.81 88.82 86.75 89.03 96.33 93.72 98.75 注:*表示结果引自相应文献,仅供参考;“—”表示缺少同环境下的可比效率结果。前向耗时为单幅图像的模型前向平均耗时,不含图像读取、预处理、后处理、结果保存及指标计算;帧率据此前向耗时换算。 表 5 目标检测误差分析结果
数据集 模型 漏检率/% 平均误检数/幅 平均IoU/% 中心偏移误差 尺度误差/% IRSTD-1k RT-DETR-R18 15.70 0.490 74.84 0.0575 10.73 IRSTD-1k FSP-DETR 13.22 0.380 75.15 0.0520 10.45 NUAA-SIRST RT-DETR-R18 10.71 0.209 76.44 0.0533 10.63 NUAA-SIRST FSP-DETR 5.36 0.070 76.80 0.0491 11.05 -
[1] ZHANG Nan, LIU Youmeng, LIU Hao, et al. DTNet: A specialized dual-tuning network for infrared vehicle detection in aerial images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5002815. doi: 10.1109/TGRS.2024.3386309. [2] LI Boyang, XIAO Chao, WANG Longguang, et al. Dense nested attention network for infrared small target detection[J]. IEEE Transactions on Image Processing, 2023, 32: 1745–1758. doi: 10.1109/TIP.2022.3199107. [3] LIN Fanzhao, BAO Kexin, LI Yong, et al. Learning contrast-enhanced shape-biased representations for infrared small target detection[J]. IEEE Transactions on Image Processing, 2024, 33: 3047–3058. doi: 10.1109/TIP.2024.3391011. [4] LI Fenghong, RAO Peng, SUN Wen, et al. A low signal-to-noise ratio infrared small-target detection network[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 8643–8658. doi: 10.1109/JSTARS.2025.3550581. [5] LI Qiang, ZHANG Mingwei, YANG Zhigang, et al. Edge-guided perceptual network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5643510. doi: 10.1109/TGRS.2024.3471865. [6] YANG Huoren, MU Tingkui, DONG Ziyue, et al. PBT: Progressive background-aware transformer for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5004513. doi: 10.1109/TGRS.2024.3415080. [7] ZHANG Shizhou, WANG Zhang, XING Yinghui, et al. SCAFNet: Semantic-guided cascade adaptive fusion network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007712. doi: 10.1109/TGRS.2024.3492256. [8] 张晶晶, 曹思华, 崔文楠, 等. 基于改进顶帽变换的红外弱小目标检测[J]. 电子与信息学报, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562.ZHANG Jingjing, CAO Sihua, CUI Wennan, et al. Improved top-hat transform-based algorithm for infrared dim and small target detection[J]. Journal of Electronics & Information Technology, 2024, 46(1): 267–276. doi: 10.11999/JEIT221562. [9] 薛驰, 陈小梅, 李海彤. 融合背景估计与相对局部对比度的天基短波红外弱小目标检测[J]. 光学 精密工程, 2026, 34(3): 450–465. doi: 10.37188/OPE.20263403.0450.XUE Chi, CHEN Xiaomei, and LI Haitong. SWIR weak targets detection on space-based platform integrating background estimation and relative local contrast[J]. Optics and Precision Engineering, 2026, 34(3): 450–465. doi: 10.37188/OPE.20263403.0450. [10] YANG Bo, ZHANG Xinyu, ZHANG Jian, et al. EFLNet: Enhancing feature learning network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5906511. doi: 10.1109/TGRS.2024.3365677. [11] CHEN Tianxiang, TAN Zhentao, GONG Tao, et al. Feature preservation and shape cues assist infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5006412. doi: 10.1109/TGRS.2024.3461795. [12] 盛卫东, 吴双林, 肖超, 等. 可微稀疏掩模引导的红外小目标快速检测网络[J]. 电子与信息学报, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989.SHENG Weidong, WU Shuanglin, XIAO Chao, et al. Differentiable sparse mask guided infrared small target fast detection network[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4779–4789. doi: 10.11999/JEIT250989. [13] 吕鹏远, 兰金江, 曾学仁, 等. 基于特征增强与融合的红外目标检测算法[J]. 红外技术, 2024, 46(7): 782–790.LYU Pengyuan, LAN Jinjiang, ZENG Xueren, et al. Infrared object detection algorithm based on feature enhancement and fusion[J]. Infrared Technology, 2024, 46(7): 782–790. [14] YUAN Shuai, QIN Hanlin, YAN Xiang, et al. SCTransNet: Spatial-channel cross transformer network for infrared small target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5002615. doi: 10.1109/TGRS.2024.3383649. [15] CHEN Tianxiang, YE Zi, TAN Zhentao, et al. MiM-ISTD: Mamba-in-mamba for efficient infrared small-target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007613. doi: 10.1109/TGRS.2024.3485721. [16] MA Tianfeng, GUO Guanqin, LI Zhimeng, et al. Infrared small target detection method based on high-low-frequency semantic reconstruction[J]. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 6012505. doi: 10.1109/LGRS.2024.3428626. [17] XU Mingzhu, YU Chenglong, LI Zexuan, et al. HDNet: A hybrid domain network with multiscale high-frequency information enhancement for infrared small-target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5004115. doi: 10.1109/TGRS.2025.3574962. [18] 刘皓皎, 刘力双, 张明淳. 基于YOLOv5改进的红外目标检测算法[J]. 激光技术, 2024, 48(4): 534–541. doi: 10.7510/jgjs.issn.1001-3806.2024.04.011.LIU Haojiao, LIU Lishuang, and ZHANG Mingchun. An improved infrared object detection algorithm based on YOLOv5[J]. Laser Technology, 2024, 48(4): 534–541. doi: 10.7510/jgjs.issn.1001-3806.2024.04.011. [19] 汪佳旭, 杨俊, 许聪源. 基于YOLO的自适应多尺度红外目标检测网络[J]. 光电工程, 2026, 53(4): 250292. doi: 10.12086/oee.2026.250292.WANG Jiaxu, YANG Jun, and XU Congyuan. YOLO-based adaptive multi-scale infrared target detection network[J]. Opto-Electronic Engineering, 2026, 53(4): 250292. doi: 10.12086/oee.2026.250292. [20] 王坤, 丁麒龙. 利用自适应融合和混合锚检测器的遥感图像小目标检测算法[J]. 电子与信息学报, 2024, 46(7): 2942–2951. doi: 10.11999/JEIT230966.WANG Kun and DING Qilong. Remote sensing images small object detection algorithm with adaptive fusion and hybrid anchor detector[J]. Journal of Electronics & Information Technology, 2024, 46(7): 2942–2951. doi: 10.11999/JEIT230966. [21] 刘杰, 刘书豪, 田明, 等. 复杂环境下无人机航拍小目标检测算法[J]. 电子与信息学报, 2026, 48(4): 1763–1773. doi: 10.11999/JEIT251126.LIU Jie, LIU Shuhao, TIAN Ming, et al. Small object detection algorithm for UAV aerial images in complex environments[J]. Journal of Electronics & Information Technology, 2026, 48(4): 1763–1773. doi: 10.11999/JEIT251126. [22] XIAO Bin and HU Yue. Bounding box regression network for infrared small target detection with adaptive receptive field and cross-scale fusion[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5003514. doi: 10.1109/TGRS.2025.3564958. [23] 张上, 崔玉杰, 龚国强, 等. LIST-DETR: 复杂背景下红外小目标检测的高效轻量级检测网络[J]. 红外与激光工程, 2025, 54(10): 20250271. doi: 10.3788/IRLA20250271.ZHANG Shang, CUI Yujie, GONG Guoqiang, et al. LIST-DETR: An efficient and lightweight detection network for infrared small target detection in complex backgrounds[J]. Infrared and Laser Engineering, 2025, 54(10): 20250271. doi: 10.3788/IRLA20250271. [24] ZHAO Yian, LV Wenyu, XU Shangliang, et al. DETRs beat YOLOs on real-time object detection[C]. Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 16965–16974. doi: 10.1109/CVPR52733.2024.01605. -
下载: