Multi-frequency Feature Interaction and Dynamic Fusion for Spacecraft 6D Pose Estimation
-
摘要: 航天器的6自由度位姿估计,是指确定目标航天器相对于服务航天器在空间坐标系中的相对姿态的处理过程。该技术是失效卫星清理、太空垃圾捕获、航天器在轨操控、空间站交会对接等一系列近距离操作任务的关键步骤。近年来,基于卷积神经网络(CNN)的6D位姿估计方法受到广泛关注,但其过度依赖卷积网络结构,容易导致对图像纹理的敏感,并缺乏对远程上下文信息的有效建模能力。此外,当前主流方法通常采用由目标检测与姿态估计构成的流水线设计,其特征提取多样性有限,对低光照和复杂背景干扰较为敏感。针对上述问题,该文提出一种基于多频谱特征交互与动态融合的航天器位姿估计网络。首先,该网络利用不同感受野的主干网络分别提取图像的空间高频特征(如语义与边缘细节)和空间低频特征(如全局结构信息);在此基础上,基于Transformer的特征匹配机制,进行自特征与跨特征注意力交互,实现长范围上下文信息的聚合;为进一步利用丰富的频域特征表示,引入频率引导特征模块,动态融合多频谱特征。最后,在航天器位姿估计数据集上的大量实验表明,所提方法具有竞争力,并展现了较强的泛化能力,相比于现有方法具有优势。Abstract:
Spacecraft Six-Dimensional (6D) pose estimation aims to determine the relative pose of a target spacecraft with respect to a servicing spacecraft in a spatial coordinate system. It is a key step in close-range operations, including defunct satellite removal, space debris capture, on-orbit spacecraft manipulation, and space station rendezvous and docking. In recent years, Convolutional Neural Network (CNN)-based 6D pose estimation methods have received considerable attention. However, their strong dependence on convolutional architectures makes them sensitive to image textures and limits their ability to model long-range contextual information effectively. Moreover, mainstream methods typically adopt a pipeline consisting of object detection followed by pose estimation. Such methods provide limited diversity in feature extraction and are sensitive to low-light conditions and complex background interference. To address these problems, a spacecraft pose estimation network based on multi-frequency feature interaction and dynamic fusion is proposed. Backbone networks with different receptive fields are first used to separately extract spatial high-frequency features, such as semantic and edge details, and spatial low-frequency features, such as global structural information. A Transformer-based feature interaction mechanism is then used to perform self-attention and cross-attention between the two types of features, thereby aggregating long-range contextual information. To further exploit rich frequency-domain representations, a frequency-guided feature module is introduced to dynamically fuse multi-frequency features. Extensive experiments on spacecraft pose estimation datasets demonstrate that the proposed method achieves competitive performance and strong generalization capability compared with existing methods. Objective The local receptive fields of CNN operators limit their ability to capture long-range dependencies. By contrast, Transformers with multi-head self-attention are more effective in globally modeling long-range contextual information and are well suited to multiscale targets and small spacecraft at long distances. Therefore, improving the extraction and interaction of spatial high- and low-frequency features is essential for increasing the accuracy of spacecraft pose estimation. Methods A 6D pose estimation network based on Transformer-based multi-frequency feature interaction and dynamic fusion is proposed. The framework consists of two branches. A ResNet-18 branch is used to extract spatial high-frequency features, including semantic and edge information, which are also used for spacecraft semantic segmentation. A DarkNet-53 branch is used to extract spatial low-frequency features that emphasize global structural information. A Transformer-based feature interaction module is then designed to model interactions between the high- and low-frequency features through self-attention and cross-attention. This mechanism aggregates long-range contextual information, increases the diversity of feature representations for texture-poor spacecraft images, and facilitates network optimization. A frequency-guided dynamic feature fusion module is further introduced to integrate the multi-frequency features. Results and Discussions The proposed method estimates spacecraft 6D pose through multi-frequency feature interaction and dynamic fusion. On the SwissCube dataset, it achieves an overall ADI-0.1d accuracy of 82.31%, exceeding those of CA-SpaceNet (79.39%) and WDR* (78.78%). On the SPEED dataset, the combined error eq + et is 0.0299, which is 22.3% lower than that of CA-SpaceNet (0.0385) and 25.2% lower than that of WDR* (0.0400). The qualitative results in Fig. 5 andFig. 7 show that the predicted keypoints are closer to the ground truth, particularly for challenging images dominated by low-frequency content and affected by complex backgrounds. Ablation experiments further show that the feature interaction module alone increases the overall accuracy from 79.39% to 80.67%, whereas the addition of frequency-guided dynamic fusion further increases it to 82.31%. These results demonstrate that the proposed framework improves the robustness of spacecraft pose estimation and provides more reliable pose predictions for on-orbit servicing and spacecraft rendezvous and docking.Conclusions To improve the accuracy of spacecraft 6D pose estimation under challenging space imaging conditions, a multi-frequency feature interaction mechanism and a semantic segmentation branch using semantic and edge information are introduced. Long-range contextual information from spatial high- and low-frequency features is aggregated to improve feature representation. In addition, a frequency-guided dynamic feature fusion module is introduced to exploit rich frequency-domain representations. Comprehensive comparisons with mainstream methods on the SwissCube and SPEED datasets demonstrate that the proposed method improves feature representation and increases the robustness of spacecraft pose estimation under complex interference. -
表 1 与主流方法在SWISSCUBE数据集的定量比较结果
方法 Near ↑ Medium ↑ Far ↑ All ↑ SegDriven[26] 41.1 22.9 7.1 21.8 SegDriven-Z[26] 52.6 45.4 29.4 43.2 DLR[17] 63.8 47.8 28.9 46.8 WDR 65.2 48.7 31.9 47.9 WDR* 92.37 84.16 61.27 78.78 CA-SpaceNet 91.01 86.32 61.72 79.39 DTSE-SpaceNet 92.57 88.74 64.42 81.65 WDR +本文方法 96.24 88.82 63.13 81.09 CA-SpaceNet+本文方法 95.76 90.61 65.12 82.31 表 2 在SPEED数据集的交叉验证定量比较结果
指标 1 2 3 4 5 mean std Mean $ {e}_{{\mathrm{q}}} $ 0.023058 0.022846 0.022311 0.023072 0.022629 0.022783 0.000286 Median $ {e}_{{\mathrm{q}}} $ 0.017626 0.017660 0.017817 0.017887 0.017896 0.017772 0.000113 Mean $ {e}_{{\mathrm{t}}} $ 0.007490 0.007149 0.007460 0.007267 0.007178 0.007308 0.000141 Median $ {e}_{{\mathrm{t}}} $ 0.005309 0.005074 0.004966 0.005180 0.005117 0.005129 0.000114 Mean S 0.030548 0.029995 0.029771 0.030339 0.029807 0.030092 0.000304 Median S 0.022935 0.022734 0.022783 0.023067 0.023013 0.0229064 0.000128 表 3 与主流方法在SPEED数据集的定量比较结果
表 4 不同模型组合方法在SWISSCUBE数据集的定量比较结果
方法 Near ↑ Medium ↑ Far ↑ All ↑ CA-SpaceNet 91.01 86.32 61.72 79.39 CA-SpaceNet +
Transformer特征交互95.02 88.20 63.49 80.67 (+4.01) (+1.88) (+1.77) (+1.28) 本文方法 95.76 90.61 65.12 82.31 (+4.65) (+4.29) (+3.40) (+2.92) 表 5 与主流方法的模型参数定量比较结果
模型 模型参数(M) 模型大小(MB) WDR* 52.1 205.2 CA-SpaceNet 51.3 205.17 DTSE-SpaceNet 51.9 206.2 本文方法 62.4 249.6 -
[1] LIU Yating, QI Zhaoshuai, CHEN Pulin, et al. TAP-Track: Generalizable spacecraft pose tracking by tracking any points[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5630713. doi: 10.1109/TGRS.2025.3584925. [2] JIANG Cuicui, GUO Pengyu, HU Qinglei, et al. Uncooperative spacecraft pose estimation based on intensity and range images fusion[J]. IEEE Transactions on Instrumentation and Measurement, 2024, 73: 5028910. doi: 10.1109/TIM.2024.3441021. [3] ZHONG Lijun, CHEN Shengpeng, WANG Wei, et al. Uncooperative spacecraft pose estimation with normalized segmentation coordinate space[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(3): 2293–2304. doi: 10.1109/TMECH.2024.3442570. [4] LIU Zibin, GUAN Banglei, SHANG Yang, et al. Stereo event-based, 6-DOF pose tracking for uncooperative spacecraft[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5607513. doi: 10.1109/TGRS.2025.3530915. [5] GUO Pengyu, JIANG Cuicui, LONG Chenrong, et al. Noncooperative spacecraft pose measurement without prior knowledge based on SAM2[J]. IEEE Transactions on Instrumentation and Measurement, 2026, 75: 5001311. doi: 10.1109/TIM.2026.3654753. [6] HAN Bing, WANG Chenxi, ZHANG Xinyu, et al. Pose estimation and neural implicit reconstruction toward noncooperative spacecraft without offline prior information[J]. IEEE Transactions on Aerospace and Electronic Systems, 2025, 61(2): 2612–2630. doi: 10.1109/TAES.2024.3479199. [7] 张春云, 孟昕曈, 陶陶, 等. 面向机器人螺栓装配的视觉感知与力控协同方法[J]. 电子与信息学报, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.ZHANG Chunyun, MENG Xintong, TAO Tao, et al. Vision-guided and force-controlled method for robotic screw assembly[J]. Journal of Electronics & Information Technology, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193. [8] 童伟, 张苗苗, 李东方, 等. 基于边缘辅助极线Transformer的多视角场景重建[J]. 电子与信息学报, 2023, 45(10): 3483–3491. doi: 10.11999/JEIT221244.TONG Wei, ZHANG Miaomiao, LI Dongfang, et al. Multiview scene reconstruction based on edge assisted epipolar Transformer[J]. Journal of Electronics & Information Technology, 2023, 45(10): 3483–3491. doi: 10.11999/JEIT221244. [9] 陈丹, 陈浩, 王子晨, 等. 多层ICP闭环检测下的误差状态卡尔曼滤波多模态融合SLAM[J]. 电子与信息学报, 2025, 47(5): 1517–1528. doi: 10.11999/JEIT240980.CHEN Dan, CHEN Hao, WANG Zichen, et al. Error state Kalman filter multimodal fusion SLAM based on MICP closed-loop detection[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1517–1528. doi: 10.11999/JEIT240980. [10] WANG Jinghao, LI Zhang, SUN Cong, et al. Satellite pose set estimation by uncertainty-guided conformal keypoint detection[J]. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36(12): 20120–20132. doi: 10.1109/TNNLS.2025.3598481. [11] PROENÇA P F and GAO Yang. Deep learning for spacecraft pose estimation from photorealistic rendering[C]. 2020 IEEE International Conference on Robotics and Automation, Paris, France, 2020: 6007–6013. doi: 10.1109/ICRA40945.2020.9197244. [12] SUN Han, ZHOU Zhenning, WANG Yizhao, et al. FGCT6D: Frequency-guided CNN-Transformer fusion network for metal parts’ robust 6D pose estimation[J]. IEEE Robotics and Automation Letters, 2024, 9(5): 4385–4392. doi: 10.1109/LRA.2024.3381016. [13] PAVLAKOS G, ZHOU Xiaowei, CHAN A, et al. 6-DoF object pose from semantic keypoints[C]. IEEE International Conference on Robotics and Automation, Singapore, Singapore, 2017: 2011–2018. doi: 10.1109/ICRA.2017.7989233. [14] SHARMA S, VENTURA J, and D’AMICO S. Robust model-based monocular pose initialization for noncooperative spacecraft rendezvous[J]. Journal of Spacecraft and Rockets, 2018, 55(6): 1414–1429. doi: 10.2514/1.A34124. [15] AUGENSTEIN S and ROCK S M. Improved frame-to-frame pose tracking during vision-only SLAM/SFM with a tumbling target[C]. 2011 IEEE International Conference on Robotics and Automation, Shanghai, China, 2011: 3131–3138. doi: 10.1109/ICRA.2011.5980232. [16] PARK T H, SHARMA S, and D'AMICO S. Towards robust learning-based pose estimation of noncooperative spacecraft[EB/OL]. https://doi.org/10.48550/arXiv.1909.00392, 2019. [17] CHEN Bo, GAO Jiewei, PARRA A, et al. Satellite pose estimation with deep landmark regression and nonlinear pose refinement[C]. 2019 IEEE/CVF International Conference on Computer Vision Workshop, Seoul, Korea (South), 2019: 2816–2824. doi: 10.1109/ICCVW.2019.00343. [18] SUN Ke, XIAO Bin, LIU Dong, et al. Deep high-resolution representation learning for human pose estimation[C]. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 5686–5696. doi: 10.1109/CVPR.2019.00584. [19] HU Yinlin, SPEIERER S, JAKOB W, et al. Wide-depth-range 6D object pose estimation in space[C]. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 15865–15874. doi: 10.1109/CVPR46437.2021.01561. [20] WANG Shunli, WANG Shuaibing, JIAO Bo, et al. Counterfactual analysis for 6D pose estimation in space[C]. 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems, Kyoto, Japan, 2022: 10627–10634. doi: 10.1109/IROS47612.2022.9981172. [21] WANG Zi, ZHANG Zhuo, SUN Xiaoliang, et al. Revisiting monocular satellite pose estimation with transformer[J]. IEEE Transactions on Aerospace and Electronic Systems, 2022, 58(5): 4279–4294. doi: 10.1109/TAES.2022.3161605. [22] CARION N, MASSA F, SYNNAEVE G, et al. End-to-end object detection with transformers[C]. The 16th European Conference on Computer Vision, Glasgow, UK, 2020: 213–229. doi: 10.1007/978-3-030-58452-8_13. [23] LIU Fengyi, ZHANG Zhujun, and LI Sijue. DTSE-SpaceNet: Deformable-transformer-based single-stage end-to-end network for 6-D pose estimation in space[J]. IEEE Transactions on Aerospace and Electronic Systems, 2024, 60(3): 2555–2571. doi: 10.1109/TAES.2023.3332075. [24] DING Yikang, YUAN Wentao, ZHU Qingtian, et al. TransMVSNet: Global context-aware multi-view stereo network with Transformers[C]. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 8575–8584. doi: 10.1109/CVPR52688.2022.00839. [25] QIN Zequn, ZHANG Pengyi, WU Fei, et al. FcaNet: Frequency channel attention networks[C]. 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 763–772. doi: 10.1109/ICCV48922.2021.00082. [26] HU Yinlin, HUGONOT J, FUA P, et al. Segmentation-driven 6D object pose estimation[C]. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 3380–3389. doi: 10.1109/CVPR.2019.00350. [27] KISANTAL M, SHARMA S, PARK T H, et al. Satellite pose estimation challenge: Dataset, competition design, and results[J]. IEEE Transactions on Aerospace and Electronic Systems, 2020, 56(5): 4083–4098. doi: 10.1109/TAES.2020.2989063. -
下载: