Multi-Agent Deep Reinforcement Learning Strategy for Multi-Spacecraft Long-Distance Orbital Pursuit-Evasion Games
-
摘要: 针对多航天器轨道追逃博弈问题,该文构建了一个尚未得到系统研究的新型研究场景。为提升航天器决策能力,使其在复杂多智能体博弈环境下生成更加稳健的博弈策略,文中提出一种基于渐进式对抗训练框架的多智能体深度强化学习算法,用于求解各航天器博弈策略。研究设置两组具备不同轨道特性的算例以及多组仿真条件开展数值模拟,并结合行为偏差分析完成策略鲁棒性验证,探究轨道特性、仿真条件与行为偏差对航天器博弈策略的影响规律。仿真结果表明,所提算法能够为各航天器生成有效博弈策略,所得策略满足全部约束条件,具备良好的鲁棒性能。
-
关键词:
- 多航天器追逃博弈 /
- 多智能体深度强化学习 /
- 渐进式对抗训练框架 /
- 行为偏差分析
Abstract:This paper presents a novel research scenario for the multi-spacecraft Orbital Pursuit-Evasion Game (OPEG), which has not yet been systematically studied. To improve spacecraft decision-making and enable robust policies in complex multi-agent games, a Multi-Agent Deep Reinforcement Learning (MADRL) algorithm based on a Progressive Adversarial Training Framework (PATF) is proposed to solve the game policies of each spacecraft. Two numerical cases with different orbital characteristics and four simulation setups are designed for verification. Behavioral deviation analysis is also conducted to evaluate policy robustness. The effects of different orbital characteristics, simulation setups, and behavioral deviations on spacecraft game policies are analyzed. Simulation results show that the proposed method enables each spacecraft to develop effective game policies that satisfy all prescribed constraints and show good robustness. Objective As the space environment becomes increasingly complex, space security has become a major research topic. A large amount of space debris and failed spacecraft pose serious threats to high-value spacecraft in orbit. Therefore, the OPEG for non-cooperative target spacecraft has attracted considerable attention. Existing studies mainly focus on two-spacecraft OPEGs, whereas multi-spacecraft OPEGs remain less explored. When more than two participants are included, a zero-sum game formulation is no longer feasible, and the problem becomes difficult to solve using traditional methods. Moreover, existing studies often ignore engineering dynamic constraints and simplify the dynamics or define the problem in a two-dimensional scene, which may introduce considerable errors. To address these limitations, this paper proposes a novel multi-spacecraft OPEG scenario. The aim is to investigate the application of MADRL to solving the approximate steady-state policies of each spacecraft in a long-distance multi-spacecraft OPEG. This study highlights the advantages of MADRL in multi-spacecraft OPEGs and provides a feasible approach for future autonomous multi-spacecraft game decision-making. Methods A Multi-Agent Proximal Policy Optimization (MAPPO) algorithm based on PATF is used to solve the approximate steady-state policy for each spacecraft in the multi-spacecraft OPEG. First, a multi-constrained multi-spacecraft OPEG model is established based on practical engineering constraints, and the problem is formulated as a Partially Observable Stochastic Game (POSG). Second, to improve the decision-making ability of agents in complex multi-agent game environments and develop more robust game policies, a novel PATF is proposed. Different reward functions are designed for the specific missions of each spacecraft. Finally, two numerical cases with different orbital characteristics are designed. Four simulation setups are then used for simulation and behavioral deviation analysis. Results and Discussions The proposed PATF-based MAPPO algorithm is compared with the original MAPPO algorithm ( Fig. 3 ). The results show that the proposed method learns effective policies more rapidly, reduces ineffective exploration, and achieves a higher final convergence reward with smaller fluctuations in the reward curve. These results also show that PATF can significantly improve the decision-making ability of agents and help them develop robust policies more effectively. Simulation verification is conducted using two numerical cases under four different setups (Figs. 4 ,5 ,6 , and7 ). The simulation results (Tables 3 and4 ) show that the proposed method performs well in both cases. The results further show that the pursuer is more likely to be intercepted when the pursuer and interceptor are in the same orbital plane. When the interceptor and the target are not in the same orbital plane, the interception mission becomes relatively easier. This paper also analyzes behavioral deviations on both sides of the game by adding control noise. The simulation results (Tables 5 and6 ) show that both sides adopt relatively conservative policies to counter control noise. The game policy obtained by the proposed method is an approximate steady-state policy. Behavioral deviations reduce the deviating side’s payoff and increase the opponent’s payoff, while the game policy maintains good robustness.Conclusions The proposed method can be effectively applied to long-distance OPEGs with multiple spacecraft in non-coplanar elliptical orbits, enabling each spacecraft to develop effective game policies. PATF improves spacecraft decision-making in complex multi-spacecraft dynamic systems, and robust control policies are developed by both the pursuer and the interceptors. The results also demonstrate the accuracy and effectiveness of the reward function design. Based on two numerical cases and simulation results under different setups, the effects of different orbital characteristics on the policies of both sides are analyzed. When the interceptors have different maximum thrusts, the decision-making of each spacecraft changes accordingly. Behavioral deviation analysis shows that the game policies of each spacecraft have good robustness. When one side’s behavior deviates, the approximate steady-state policy balance changes, which reduces its own payoff and increases the opponent’s payoff. The research scenario proposed in this paper expands the scope of existing studies on multi-spacecraft game problems. -
表 1 算例1和算例2中航天器的轨道根数
轨道根数 算例 1 算例 2 追击者 拦截者1 拦截者2 追击者 拦截者1 拦截者2 $ a $(km) 16000 16200 16200 15900 16200 16200 $ e $ 0.2 0.2 0.2 0.2 0.2 0.2 $ i $ (°) 35 36 35 35.5 35.5 35 $ {\varOmega } $ (°) 30 30 30 30 30 30 $ \omega $ (°) 91 89 90 89 89 89.7 $ f $ (°) 0 0 0 0 0.735 0 表 2 多航天器OPEG模型的参数设置
参数 $ N $ $ \Delta t $ $ {T}_{\mathrm{P},\max } $ $ {T}_{\mathrm{I},\max } $ $ \Delta {r}_{\mathrm{P},\max } $ $ \Delta {r}_{\mathrm{td},\max } $ $ \Delta {r}_{\mathrm{I},\max } $ $ \Delta {r}_{\mathrm{sd},\min } $ $ \Delta {r}_{\mathrm{sd},\max } $ 值 50 60 s 100 N 60 N, 75 N 3 km 20 km 5 km 2 km 5 km 表 3 算例1的数值结果
设置 追击者 拦截者1 拦截者2 步数 胜利者 $ \Delta {r}_{\text{PT}} $ (km) $ \Delta {r}_{\text{PI}} $ (km) $ {\Delta m}_{\mathrm{P}} $ (kg) $ \Delta {r}_{{{\mathrm{I}}_{1}}\mathrm{P}} $ (km) $ \Delta {r}_{{{\mathrm{I}}_{1}}\mathrm{T}} $ (km) $ \Delta {m}_{{{\mathrm{I}}_{1}}} $ (kg) $ \Delta {r}_{{{\mathrm{I}}_{2}}\mathrm{P}} $ (km) $ \Delta {r}_{{{\mathrm{I}}_{2}}\mathrm{T}} $ (km) $ {\Delta m}_{{{\mathrm{I}}_{2}}} $ (kg) 1 1.23 16.32 15.85 201.79 202.94 14.86 16.32 16.38 17.38 38 P 2 41.42 0.74 17.44 235.39 268.26 14.54 0.74 41.21 17.27 32 $ {\mathrm{I}}_{2} $ 3 1.12 12.13 22.44 106.07 105.73 21.35 12.13 12.49 15.99 38 P 4 41.41 0.76 17.45 183.42 218.56 18.22 0.76 41.41 17.26 32 $ {\mathrm{I}}_{2} $ 表 4 算例2的数值结果
设置 追击者 拦截者1 拦截者2 步数 胜利者 $ \Delta {r}_{\text{PT}} $ (km) $ \Delta {r}_{\text{PI}} $ (km) $ {\Delta m}_{\mathrm{P}} $ (kg) $ \Delta {r}_{{{\mathrm{I}}_{1}}\mathrm{P}} $ (km) $ \Delta {r}_{{{\mathrm{I}}_{1}}\mathrm{T}} $ (km) $ \Delta {m}_{{{\mathrm{I}}_{1}}} $ (kg) $ \Delta {r}_{{{\mathrm{I}}_{2}}\mathrm{P}} $ (km) $ \Delta {r}_{{{\mathrm{I}}_{2}}\mathrm{T}} $ (km) $ {\Delta m}_{{{\mathrm{I}}_{2}}} $ (kg) 1 0.14 26.15 10.98 35.63 35.72 18.32 26.15 26.20 18.32 40 P 2 0.26 22.24 13.62 55.58 55.49 16.48 22.24 22.37 17.50 36 P 3 76.01 0.49 18.24 0.49 75.84 19.81 103.14 82.06 14.09 36 $ {\mathrm{I}}_{1} $ 4 51.66 1.15 18.64 1.15 52.23 20.29 31.62 34.70 20.40 36 $ {\mathrm{I}}_{1} $ 表 5 追击者具有控制噪声的仿真数值结果
设置 $ {\Delta m}_{\mathrm{P}} $ (kg) $ \Delta {m}_{{{\mathrm{I}}_{1}}} $ (kg) $ {\Delta m}_{{{\mathrm{I}}_{2}}} $ (kg) 步数 MSR (%) 胜利者 mean std mean std mean std mean std 1 16.59 0.55 15.44 0.089 16.65 0.16 39.17 1.34 98.5 P 2 15.20 0.051 12.55 9.7×10-5 14.86 1.8×10-4 31 0 100 I 3 21.09 1.34 21.55 0.56 15.16 0.85 41.41 1.87 95.5 P 4 15.35 0.058 14.93 2.4×10-4 15.00 6.7×10-5 29 0 100 I 表 6 拦截者具有控制噪声的仿真数值结果
设置 $ {\Delta m}_{\mathrm{P}} $ (kg) $ \Delta {m}_{{{\mathrm{I}}_{1}}} $ (kg) $ {\Delta m}_{{{\mathrm{I}}_{2}}} $ (kg) 步数 MSR (%) 胜利者 mean std mean std mean std mean std 1 14.02 7.8×10-5 14.50 0.043 14.84 0.040 39 0 100 P 2 19.98 9.7×10-5 16.72 0.051 18.43 0.037 40 0 100 P 3 18.77 1.2×10-4 18.42 0.059 15.43 0.051 38 0 100 P 4 22.03 7.2×10-5 19.91 0.063 20.89 0.034 38 0 100 P -
[1] SUN Qilong, QI Naiming, XIAO Longxu, et al. Differential game strategy in three-player evasion and pursuit scenarios[J]. Journal of Systems Engineering and Electronics, 2018, 29(2): 352–366. doi: 10.21629/JSEE.2018.02.16. [2] CUI Jianfeng, LI Dongchang, LIU Peng, et al. Game-model prediction hybrid path planning algorithm for multiple mobile robots in pursuit evasion game[C]. 2021 IEEE International Conference on Unmanned Systems (ICUS), Beijing, China, 2021: 925–930. doi: 10.1109/ICUS52573.2021.9641362. [3] ZHANG Yiqun, ZHANG Pengfei, WANG Xiaodong, et al. An open loop Stackelberg solution to optimal strategy for UAV pursuit-evasion game[J]. Aerospace Science and Technology, 2022, 129: 107840. doi: 10.1016/j.ast.2022.107840. [4] 高思华, 刘宝煜, 惠康华, 等. 信息年龄约束下的无人机数据采集能耗优化路径规划算法[J]. 电子与信息学报, 2024, 46(10): 4024–4034. doi: 10.11999/JEIT240075.GAO Sihua, LIU Baoyu, HUI Kanghua, et al. Energy-efficient UAV trajectory planning algorithm for AoI-constrained data collection[J]. Journal of Electronics & Information Technology, 2024, 46(10): 4024–4034. doi: 10.11999/JEIT240075. [5] 颜志, 陆元媛, 丁聪, 等. 面向用户移动场景的无人机中继功率分配与轨迹设计[J]. 电子与信息学报, 2024, 46(5): 1896–1907. doi: 10.11999/JEIT231337.YAN Zhi, LU Yuanyuan, DING Cong, et al. Power allocation and trajectory design for unmanned aerial vehicle relay network with mobile users[J]. Journal of Electronics & Information Technology, 2024, 46(5): 1896–1907. doi: 10.11999/JEIT231337. [6] LOWE R, WU Yi, TAMAR A, et al. Multi-agent actor-critic for mixed cooperative-competitive environments[C]. The 31st International Conference on Neural Information Processing Systems, Long Beach California, USA, 2017: 6382–6393. [7] LUO Yuelin, GANG Tieqiang, and CHEN Lijie. Research on target defense strategy based on deep reinforcement learning[J]. IEEE Access, 2022, 10: 82329–82335. doi: 10.1109/ACCESS.2022.3179373. [8] WANG Xin, WANG Yueying, ZHOU Weixiang, et al. Pursuit-evasion game of unmanded surface vehicles based on deep reinforcement learning[C]. 2023 4th International Conference on Electronic Communication and Artificial Intelligence (ICECAI), Guangzhou, China, 2023: 358–363. doi: 10.1109/ICECAI58670.2023.10176487. [9] JI Mengda, XU Genjiu, DUAN Zekun, et al. Cooperative pursuit with multiple pursuers based on deep minimax Q-learning[J]. Aerospace Science and Technology, 2024, 146: 108919. doi: 10.1016/j.ast.2024.108919. [10] JAGAT A and SINCLAIR A J. Nonlinear control for spacecraft pursuit-evasion game using the state-dependent Riccati equation method[J]. IEEE Transactions on Aerospace and Electronic Systems, 2017, 53(6): 3032–3042. doi: 10.1109/TAES.2017.2725498. [11] MA Huidong and ZHANG Gang. Delta-V analysis for impulsive orbital pursuit-evasion based on reachable domain coverage[J]. Aerospace Science and Technology, 2024, 150: 109243. doi: 10.1016/j.ast.2024.109243. [12] SHI Mingming, YE Dong, SUN Zhaowei, et al. Spacecraft orbital pursuit–evasion games with J2 perturbations and direction-constrained thrust[J]. Acta Astronautica, 2023, 202: 139–150. doi: 10.1016/j.actaastro.2022.10.004. [13] ZHANG Jingrui, ZHANG Kunpeng, ZHANG Yao, et al. Near-optimal interception strategy for orbital pursuit-evasion using deep reinforcement learning[J]. Acta Astronautica, 2022, 198: 9–25. doi: 10.1016/j.actaastro.2022.05.057. [14] ZHAO Liran, ZHANG Yulin, and DANG Zhaohui. PRD-MADDPG: An efficient learning-based algorithm for orbital pursuit-evasion game with impulsive maneuvers[J]. Advances in Space Research, 2023, 72(2): 211–230. doi: 10.1016/j.asr.2023.03.014. [15] TANG Xu, YE Dong, LOW K S, et al. Multi-spacecraft pursuit-evasion-defense strategy based on game theory for on-orbit spacecraft servicing[C]. 2023 IEEE Aerospace Conference, Big Sky, USA, 2023: 1–9. doi: 10.1109/AERO55745.2023.10115953. [16] XU Sihan, ZHAO Liran, ZHANG Weichen, et al. Delta-V-based cooperative strategies for orbital two-pursuer one-evader pursuit–evasion games[J]. Space: Science & Technology, 2025, 5: 0222. doi: 10.34133/space.0222. [17] LIANG Haizhao, WANG Jianying, LIU Jiaqi, et al. Guidance strategies for interceptor against active defense spacecraft in two-on-two engagement[J]. Aerospace Science and Technology, 2020, 96: 105529. doi: 10.1016/j.ast.2019.105529. [18] DI Peng, YAO Ye, LIN Zheng, et al. Trajectory optimization of spacecraft autonomous far-distance rapid rendezvous based on deep reinforcement learning[J]. Advances in Space Research, 2025, 75(1): 790–806. doi: 10.1016/j.asr.2024.09.066. [19] 魏普远, 何磊. 基于深度强化学习的自适应大邻域搜索算法在成像卫星调度问题中的应用[J]. 电子与信息学报, 2025, 47(12): 5005–5015. doi: 10.11999/JEIT251009.WEI Puyuan and HE Lei. A deep reinforcement learning enhanced adaptive large neighborhood search for imaging satellite scheduling[J]. Journal of Electronics & Information Technology, 2025, 47(12): 5005–5015. doi: 10.11999/JEIT251009. [20] MU Chaoxu, LIU Shuo, LU Ming, et al. Autonomous spacecraft collision avoidance with a variable number of space debris based on safe reinforcement learning[J]. Aerospace Science and Technology, 2024, 149: 109131. doi: 10.1016/j.ast.2024.109131. [21] YU Chao, VELU A, VINITSKY E, et al. The surprising effectiveness of PPO in cooperative multi-agent games[C]. The 36th International Conference on Neural Information Processing Systems, New Orleans, USA, 2022: 1787. [22] BATE R R, MUELLER D D, WHITE J E, et al. Fundamentals of Astrodynamics[M]. 2nd ed. Dover Publications, 2020. [23] HANSEN E A, BERNSTEIN D S, and ZILBERSTEIN S. Dynamic programming for partially observable stochastic games[C]. The 19th National Conference on Artifical Intelligence, San Jose, USA, 2004: 709–715. [24] SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal policy optimization algorithms[J]. arXiv preprint arXiv: 1707.06347, 2017. doi: 10.48550/arXiv.1707.06347. -
下载: