Advanced Search
Turn off MathJax
Article Contents
TONG Wei, LIN Xi, YAN Ying, LIN Jinxing, LI Tao, WU Qi. Multi-Frequency Feature Interaction and Adaptive Fusion for cooperative Spacecraft 6D Pose Estimation[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260452
Citation: TONG Wei, LIN Xi, YAN Ying, LIN Jinxing, LI Tao, WU Qi. Multi-Frequency Feature Interaction and Adaptive Fusion for cooperative Spacecraft 6D Pose Estimation[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260452

Multi-Frequency Feature Interaction and Adaptive Fusion for cooperative Spacecraft 6D Pose Estimation

doi: 10.11999/JEIT260452 cstr: 32379.14.JEIT260452
Funds:  National Natural Science Foundation of China(62405145), Natural Science Foundation of Jiangsu Province(BK20240641), Key Laboratory of Social Computing and Cognitive Intelligence(SCCI2023YB02) , Nanjing University of Information Science & Technology, Quality Assurance and Evaluation of higher Education in Jiangsu Province(2025JSETKT158)
  • Accepted Date: 2026-08-17
  • Rev Recd Date: 2026-08-17
  • Available Online: 2026-08-26
  • Spacecraft 6D pose estimation aims to determine the relative pose between the target spacecraft and the service spacecraft in the spatial coordinate system, which is a critical step for a series of close-range operation tasks, such as failed satellite cleaning, space debris capture, on-orbit spacecraft manipulation, and space station rendezvous and docking. In recent years, CNN-based 6D pose estimation methods have received widespread attention. However, their over-reliance on convolutional network architectures makes them sensitive to image textures and limits their capacity to effectively model long-range contextual information. Moreover, current mainstream methods typically adopt a pipeline design consisting of object detection followed by pose estimation, which suffers from limited diversity in feature extraction and is sensitive to low-light conditions and complex background interference. To address these issues, this paper proposes a spacecraft pose estimation network based on multi-spectral feature interaction and dynamic fusion. Specifically, the network first leverages backbone networks with different receptive fields to separately extract spatial high-frequency features (such as semantic and edge details) and spatial low-frequency features (such as global structural information). On this basis, a Transformer-based feature matching mechanism is employed to perform self-attention and cross-attention feature interactions, thereby aggregating long-range contextual information. To further exploit the rich frequency-domain feature representations, a frequency-guided feature module is introduced to dynamically fuse multi-spectral features. Finally, extensive experiments on spacecraft pose estimation datasets demonstrate that the proposed method achieves competitive performance and strong generalization ability, showing advantages over existing methods.  Objective  Due to the limitation of local receptive field of CNN convolution operator, the mining of remote correlation information is often not ideal. In contrast, Transformer with multi-head intra-attention mechanism is more effective in globally modeling long-range context information and is good at processing multi-scale objects and small spacecraft at a long distance. Therefore, enhancing the processing ability of low-frequency and high-frequency features is of great significance for improving the accuracy.  Methods  This work proposes a 6D pose estimation network with Transformer-based multi-frequency feature interaction and fusion. The framework consists of two parts. One is to extract high-frequency features such as image semantics and edges via ResNet18 and leverage them for spacecraft semantic segmentation directly. The other is to extract global low-frequency features such as image texture through DarkNet53. Then Transformer-based feature interaction module is designed to perform attention between high-frequency and low-frequency features, enabling the aggregation of long-range contextual information, which can enhance the diversity of texture-less spacecraft image features and promote network optimization.  Results and Discussions  The proposed method estimates spacecraft 6D pose via multi-frequency feature interaction and adaptive fusion. On the SwissCube dataset, it achieves an overall ADI-0.1d accuracy of 82.31%, outperforming CA-SpaceNet (79.39%) and WDR* (78.78%). On the SPEED dataset, the combined error eq+et is 0.0299, which is 22.3% lower than CA-SpaceNet (0.0385) and 25.2% lower than WDR* (0.0400). Qualitative results in Fig. 5 and Fig. 7 show that predicted key points are significantly closer to the ground truth, especially under challenging low-frequency conditions. Ablation studies confirm that the feature interaction module alone increases overall accuracy from 79.39% to 80.67%, and adding frequency-guided fusion further raises it to 82.31%. These results demonstrate that the proposed framework enhances pose estimation robustness and meets the stringent requirements for on-orbit servicing and space rendezvous.  Conclusions  To enhance the accuracy of spacecraft 6D pose estimation under extreme atmospheric environments, this work innovatively designs a sub-branch based on multi-frequency feature interaction and additional semantic edge segmentation, and overcomes the limitation of low feature extraction efficiency of existing methods by aggregating long-range multi-frequency features. In addition, a frequency-guided dynamic feature fusion module is introduced to fully leverage the rich frequency-domain feature representation. Comprehensive experimental comparison with mainstream methods on SwissCube and SPEED datasets verifies that the proposed work can enhance the representation of feature information and improve the robustness of spacecraft pose estimation.
  • loading
  • [1]
    LIU Yating, QI Zhaoshuai, CHEN Pulin, et al. TAP-Track: Generalizable spacecraft pose tracking by tracking any points[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5630713. doi: 10.1109/TGRS.2025.3584925.
    [2]
    JIANG Cuicui, GUO Pengyu, HU Qinglei, et al. Uncooperative spacecraft pose estimation based on intensity and range images fusion[J]. IEEE Transactions on Instrumentation and Measurement, 2024, 73: 5028910. doi: 10.1109/TIM.2024.3441021.
    [3]
    ZHONG Lijun, CHEN Shengpeng, WANG Wei, et al. Uncooperative spacecraft pose estimation with normalized segmentation coordinate space[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(3): 2293–2304. doi: 10.1109/TMECH.2024.3442570.
    [4]
    LIU Zibin, GUAN Banglei, SHANG Yang, et al. Stereo event-based, 6-DOF pose tracking for uncooperative spacecraft[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5607513. doi: 10.1109/TGRS.2025.3530915.
    [5]
    GUO Pengyu, JIANG Cuicui, LONG Chenrong, et al. Noncooperative spacecraft pose measurement without prior knowledge based on SAM2[J]. IEEE Transactions on Instrumentation and Measurement, 2026, 75: 5001311. doi: 10.1109/TIM.2026.3654753.
    [6]
    HAN Bing, WANG Chenxi, ZHANG Xinyu, et al. Pose estimation and neural implicit reconstruction toward noncooperative spacecraft without offline prior information[J]. IEEE Transactions on Aerospace and Electronic Systems, 2025, 61(2): 2612–2630. doi: 10.1109/TAES.2024.3479199.
    [7]
    张春云, 孟昕曈, 陶陶, 等. 面向机器人螺栓装配的视觉感知与力控协同方法[J]. 电子与信息学报, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.

    ZHANG Chunyun, MENG Xintong, TAO Tao, et al. Vision-guided and force-controlled method for robotic screw assembly[J]. Journal of Electronics & Information Technology, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.
    [8]
    童伟, 张苗苗, 李东方, 等. 基于边缘辅助极线Transformer的多视角场景重建[J]. 电子与信息学报, 2023, 45(10): 3483–3491. doi: 10.11999/JEIT221244.

    TONG Wei, ZHANG Miaomiao, LI Dongfang, et al. Multiview scene reconstruction based on edge assisted epipolar Transformer[J]. Journal of Electronics & Information Technology, 2023, 45(10): 3483–3491. doi: 10.11999/JEIT221244.
    [9]
    陈丹, 陈浩, 王子晨, 等. 多层ICP闭环检测下的误差状态卡尔曼滤波多模态融合SLAM[J]. 电子与信息学报, 2025, 47(5): 1517–1528. doi: 10.11999/JEIT240980.

    CHEN Dan, CHEN Hao, WANG Zichen, et al. Error state Kalman filter multimodal fusion SLAM based on MICP closed-loop detection[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1517–1528. doi: 10.11999/JEIT240980.
    [10]
    WANG Jinghao, LI Zhang, SUN Cong, et al. Satellite pose set estimation by uncertainty-guided conformal keypoint detection[J]. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36(12): 20120–20132. doi: 10.1109/TNNLS.2025.3598481.
    [11]
    PROENÇA P F and GAO Yang. Deep learning for spacecraft pose estimation from photorealistic rendering[C]. 2020 IEEE International Conference on Robotics and Automation, Paris, France, 2020: 6007–6013. doi: 10.1109/ICRA40945.2020.9197244.
    [12]
    SUN Han, ZHOU Zhenning, WANG Yizhao, et al. FGCT6D: Frequency-guided CNN-Transformer fusion network for metal parts’ robust 6D pose estimation[J]. IEEE Robotics and Automation Letters, 2024, 9(5): 4385–4392. doi: 10.1109/LRA.2024.3381016.
    [13]
    PAVLAKOS G, ZHOU Xiaowei, CHAN A, et al. 6-DoF object pose from semantic keypoints[C]. IEEE International Conference on Robotics and Automation, Singapore, Singapore, 2017: 2011–2018. doi: 10.1109/ICRA.2017.7989233.
    [14]
    SHARMA S, VENTURA J, and D’AMICO S. Robust model-based monocular pose initialization for noncooperative spacecraft rendezvous[J]. Journal of Spacecraft and Rockets, 2018, 55(6): 1414–1429. doi: 10.2514/1.A34124.
    [15]
    AUGENSTEIN S and ROCK S M. Improved frame-to-frame pose tracking during vision-only SLAM/SFM with a tumbling target[C]. 2011 IEEE International Conference on Robotics and Automation, Shanghai, China, 2011: 3131–3138. doi: 10.1109/ICRA.2011.5980232.
    [16]
    PARK T H, SHARMA S, and D'AMICO S. Towards robust learning-based pose estimation of noncooperative spacecraft[EB/OL]. https://doi.org/10.48550/arXiv.1909.00392, 2019. (查阅网上资料,未能确认文献类型,请确认).
    [17]
    CHEN Bo, GAO Jiewei, PARRA A, et al. Satellite pose estimation with deep landmark regression and nonlinear pose refinement[C]. 2019 IEEE/CVF International Conference on Computer Vision Workshop, Seoul, Korea (South), 2019: 2816–2824. doi: 10.1109/ICCVW.2019.00343.
    [18]
    SUN Ke, XIAO Bin, LIU Dong, et al. Deep high-resolution representation learning for human pose estimation[C]. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 5686–5696. doi: 10.1109/CVPR.2019.00584.
    [19]
    HU Yinlin, SPEIERER S, JAKOB W, et al. Wide-depth-range 6D object pose estimation in space[C]. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 15865–15874. doi: 10.1109/CVPR46437.2021.01561.
    [20]
    WANG Shunli, WANG Shuaibing, JIAO Bo, et al. Counterfactual analysis for 6D pose estimation in space[C]. 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems, Kyoto, Japan, 2022: 10627–10634. doi: 10.1109/IROS47612.2022.9981172.
    [21]
    WANG Zi, ZHANG Zhuo, SUN Xiaoliang, et al. Revisiting monocular satellite pose estimation with transformer[J]. IEEE Transactions on Aerospace and Electronic Systems, 2022, 58(5): 4279–4294. doi: 10.1109/TAES.2022.3161605.
    [22]
    CARION N, MASSA F, SYNNAEVE G, et al. End-to-end object detection with transformers[C]. Proceedings of the 16th European Conference on Computer Vision, Glasgow, UK, 2020: 213–229. doi: 10.1007/978-3-030-58452-8_13.
    [23]
    LIU Fengyi, ZHANG Zhujun, and LI Sijue. DTSE-SpaceNet: Deformable-transformer-based single-stage end-to-end network for 6-D pose estimation in space[J]. IEEE Transactions on Aerospace and Electronic Systems, 2024, 60(3): 2555–2571. doi: 10.1109/TAES.2023.3332075.
    [24]
    QIN Zequn, ZHANG Pengyi, WU Fei, et al. FcaNet: Frequency channel attention networks[C]. 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 763–772. doi: 10.1109/ICCV48922.2021.00082.
    [25]
    DING Yikang, YUAN Wentao, ZHU Qingtian, et al. TransMVSNet: Global context-aware multi-view stereo network with Transformers[C]. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 8575–8584. doi: 10.1109/CVPR52688.2022.00839.
    [26]
    HU Yinlin, HUGONOT J, FUA P, et al. Segmentation-driven 6D object pose estimation[C]. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 3380–3389. doi: 10.1109/CVPR.2019.00350.
    [27]
    KISANTAL M, SHARMA S, PARK T H, et al. Satellite pose estimation challenge: Dataset, competition design, and results[J]. IEEE Transactions on Aerospace and Electronic Systems, 2020, 56(5): 4083–4098. doi: 10.1109/TAES.2020.2989063.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(9)  / Tables(5)

    Article Metrics

    Article views (86) PDF downloads(14) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return