Advanced Search
Turn off MathJax
Article Contents
HUANG Jieyu, XIE Junwei, ZHANG Haowei, FENG Weike, HAN Weihang. A Reinforcement Learning Driven Power Allocation Algorithm for Collocated MIMO Radar[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260695
Citation: HUANG Jieyu, XIE Junwei, ZHANG Haowei, FENG Weike, HAN Weihang. A Reinforcement Learning Driven Power Allocation Algorithm for Collocated MIMO Radar[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260695

A Reinforcement Learning Driven Power Allocation Algorithm for Collocated MIMO Radar

doi: 10.11999/JEIT260695 cstr: 32379.14.JEIT260695
Funds:  The National Natural Science Foundation of China (62571544), The Innovation Capability Support Program of Shaanxi (2025ZC-KJXX-81), The Research Program Project of Youth Innovation Team of Shaanxi Provincial Education Department (24JP221), The Natural Science Foundation of Xi’an (26ZRKX00090)
  • Received Date: 2026-05-28
  • Accepted Date: 2026-07-03
  • Rev Recd Date: 2026-07-02
  • Available Online: 2026-07-14
  •   Objective  Traditional optimization-based power allocation algorithms for collocated MIMO radar have two fundamental limitations. First, they optimize tracking performance only for the next time step and therefore lack a full-time-horizon view of the power allocation process. This myopic strategy cannot achieve optimal multi-target tracking accuracy over extended periods, particularly when target trajectories vary substantially. Second, these algorithms rely on iterative nonlinear constrained optimization, resulting in high computational complexity. Therefore, they cannot satisfy the real-time requirements of dynamic battlefield environments where target states change rapidly. To address these limitations, this paper proposes a Reinforcement Learning (RL)-driven power allocation algorithm. Unlike conventional methods, the proposed approach formulates the power allocation problem as a Markov Decision Process (MDP) that maximizes long-term cumulative tracking accuracy. The algorithm adaptively allocates limited transmit power among multiple beams according to the current system state, balancing immediate tracking performance with long-term cumulative tracking accuracy.  Methods  The Posterior Cramér-Rao Lower Bound (PCRLB) is employed to quantify the theoretical lower bound of the tracking error for each target. The state space is constructed by combining the motion states (position and velocity) of all targets with the normalized PCRLB from the previous allocation step. The action space consists of discrete transmit power levels for each beam, subject to the total power budget and individual beam power constraints. All feasible power allocation vectors are enumerated and encoded to reduce the action-space dimensionality. The reward function is defined as the negative weighted sum of the normalized PCRLB, encouraging the agent to minimize tracking errors. The power allocation process is formulated as an MDP and solved using the Dueling Double Deep Q-Network (D3QN) algorithm. The D3QN framework incorporates three major enhancements: (1) a double-network architecture comprising an online Q-network and a target Q-network to improve training stability; (2) a dueling architecture that decomposes the Q-value into a state-value function and an action-advantage function to improve action discrimination; and (3) off-policy learning with experience replay to improve the use of historical trajectories. An ε-greedy strategy is adopted for exploration, with ε gradually decreasing during training. After offline training, the learned network directly generates real-time transmit power allocation decisions from the current system state without iterative optimization.  Results and Discussions  Simulations are conducted using three targets following the Constant Velocity (CV) model. Fixed power allocation yields the lowest tracking accuracy because of inefficient resource utilization. The traditional optimization method, which minimizes the instantaneous tracking error, achieves moderate tracking performance but remains myopic. When the discount factor $ \gamma =0 $, the D3QN algorithm achieves performance comparable to that of the traditional optimization method because both optimize only immediate rewards. In contrast, when $ \gamma =0.99 $, the D3QN algorithm significantly improves full-time-horizon tracking accuracy. The resulting power allocation strategy allocates more transmit power to distant, low-Signal-to-Noise Ratio (SNR) targets at earlier stages while reducing redundant power assigned to nearby high-SNR targets. The training curves show that $ \gamma =0.99 $ achieves a higher steady-state cumulative reward, although convergence exhibits greater oscillation because of the increased difficulty of estimating long-term returns. Furthermore, the trained D3QN network generates transmit power allocation decisions almost instantaneously, whereas the traditional optimization method must solve a constrained optimization problem at every time step, providing a substantial real-time computational advantage.  Conclusions  This paper proposes an RL-driven power allocation algorithm for collocated MIMO radar multi-target tracking that overcomes the myopic behavior and high computational complexity of conventional optimization methods. The proposed algorithm constructs the state space and reward function using the PCRLB, models the power allocation process as an MDP, and solves it using the D3QN algorithm. Simulation results demonstrate that, with an appropriate discount factor ($ \gamma =0.99 $), the proposed approach significantly improves full-time-horizon tracking accuracy. This improvement results from the agent’s ability to learn a long-term optimal policy that proactively allocates transmit power to future distant, low-SNR targets. Furthermore, the trained network enables real-time decision-making through direct forward propagation, substantially reducing computational latency compared with iterative optimization. This work provides a new approach for intelligent radar resource management in complex battlefield environments.
  • loading
  • [1]
    何子述, 程子扬, 李军, 等. 集中式MIMO雷达研究综述[J]. 雷达学报, 2022, 11(5): 805–829. doi: 10.12000/JR22128.

    HE Zishu, CHENG Ziyang, LI Jun, et al. A survey of collocated MIMO radar[J]. Journal of Radars, 2022, 11(5): 805–829. doi: 10.12000/JR22128.
    [2]
    严俊坤, 陈林, 刘宏伟, 等. 基于机会约束的MIMO雷达多波束稳健功率分配算法[J]. 电子学报, 2019, 47(6): 1230–1235. doi: 10.3969/j.issn.0372-2112.2019.06.007.

    YAN Junkun, CHEN Lin, LIU Hongwei, et al. Chance constrained based robust multibeam power allocation algorithm for MIMO radar[J]. Acta Electronica Sinica, 2019, 47(6): 1230–1235. doi: 10.3969/j.issn.0372-2112.2019.06.007.
    [3]
    李正杰, 谢军伟, 张浩为, 等. 一种低截获背景下的集中式MIMO雷达快速功率分配算法[J]. 雷达学报, 2023, 12(3): 602–615. doi: 10.12000/JR22203.

    LI Zhengjie, XIE Junwei, ZHANG Haowei, et al. A fast power allocation algorithm in a collocated MIMO radar under low interception backgrounds[J]. Journal of Radars, 2023, 12(3): 602–615. doi: 10.12000/JR22203.
    [4]
    时晨光, 丁琳涛, 汪飞, 等. 面向射频隐身的组网雷达多目标跟踪下射频辐射资源优化分配算法[J]. 电子与信息学报, 2021, 43(3): 539–546. doi: 10.11999/JEIT200636.

    SHI Chenguang, DING Lintao, WANG Fei, et al. Radio frequency stealth-based optimal radio frequency resource allocation algorithm for multiple-target tracking in radar network[J]. Journal of Electronics & Information Technology, 2021, 43(3): 539–546. doi: 10.11999/JEIT200636.
    [5]
    黄洁瑜, 张浩为, 谢军伟, 等. 一种面向隐身目标跟踪的雷达组网系统资源优化分配算法[J]. 北京航空航天大学学报, 2026, 52(2): 470–481. doi: 10.13700/j.bh.1001-5965.2023.0782.

    HUANG Jieyu, ZHANG Haowei, XIE Junwei, et al. A resource optimization allocation algorithm for radar networked system for stealth target tracking[J]. Journal of Beijing University of Aeronautics and Astronautics, 2026, 52(2): 470–481. doi: 10.13700/j.bh.1001-5965.2023.0782.
    [6]
    YUAN Ye, YI Wei, HOSEINNEZHAD R, et al. Robust power allocation for resource-aware multi-target tracking with colocated MIMO radars[J]. IEEE Transactions on Signal Processing, 2021, 69: 443–458. doi: 10.1109/TSP.2020.3047519.
    [7]
    ZHANG Haowei, SHI Junpeng, ZHANG Qiliang, et al. Antenna selection for target tracking in collocated MIMO radar[J]. IEEE Transactions on Aerospace and Electronic Systems, 2021, 57(1): 423–436. doi: 10.1109/TAES.2020.3031767.
    [8]
    丁梓航, 谢军伟, 齐铖. 基于强化学习的频控阵-多输入多输出雷达发射功率分配方法[J]. 电子与信息学报, 2023, 45(2): 550–557. doi: 10.11999/JEIT211555.

    DING Zihang, XIE Junwei, and QI Cheng. Transmit power allocation method of frequency diverse array-multi input and multi output radar based on reinforcement learning[J]. Journal of Electronics & Information Technology, 2023, 45(2): 550–557. doi: 10.11999/JEIT211555.
    [9]
    张鹏, 严俊坤, 高畅, 等. 动态电磁环境下多功能雷达一体化发射资源管理方案[J]. 雷达学报, 2025, 14(2): 456–469. doi: 10.12000/JR24230.

    ZHANG Peng, YAN Junkun, GAO Chang, et al. Integrated transmission resource management scheme for multifunctional radars in dynamic electromagnetic environments[J]. Journal of Radars, 2025, 14(2): 456–469. doi: 10.12000/JR24230.
    [10]
    LU Ziyang and GURSOY M C. Resource allocation for multi-target radar tracking via constrained deep reinforcement learning[J]. IEEE Transactions on Cognitive Communications and Networking, 2023, 9(6): 1677–1690. doi: 10.1109/TCCN.2023.3304634.
    [11]
    易伟, 袁野, 刘光宏, 等. 多雷达协同探测技术研究进展: 认知跟踪与资源调度算法[J]. 雷达学报, 2023, 12(3): 471–499. doi: 10.12000/JR23036.

    YI Wei, YUAN Ye, LIU Guanghong, et al. Recent advances in multi-radar collaborative surveillance: Cognitive tracking and resource scheduling algorithms[J]. Journal of Radars, 2023, 12(3): 471–499. doi: 10.12000/JR23036.
    [12]
    ZHANG Peng, YAN Junkun, PU Wenqiang, et al. Collaborative strategy learning based transmit resource management scheme for multiple target tracking in multi-cooperative jamming environment[J]. IEEE Transactions on Signal Processing, 2026, 74: 2053–2068. doi: 10.1109/TSP.2026.3693184.
    [13]
    陈佳美, 孙慧雯, 李玉峰, 等. 基于双深度Q网络算法的无人机辅助密集网络资源优化策略[J]. 电子与信息学报, 2025, 47(8): 2621–2629. doi: 10.11999/JEIT250021.

    CHEN Jiamei, SUN Huiwen, LI Yufeng, et al. Double deep Q network algorithm-based unmanned aerial vehicle-assisted dense network resource optimization strategy[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2621–2629. doi: 10.11999/JEIT250021.
    [14]
    周黎鸣, 余汐, 范明虎, 等. 基于双深度Q网络的多目标遥感产品生产任务调度算法[J]. 电子与信息学报, 2025, 47(8): 2819–2829. doi: 10.11999/JEIT250089.

    ZHOU Liming, YU Xi, FAN Minghu, et al. Multi-objective remote sensing product production task scheduling algorithm based on double deep Q-network[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2819–2829. doi: 10.11999/JEIT250089.
    [15]
    李一兵, 孙柳晴, 戚昌龙. 基于改进秘书鸟算法的协同干扰资源分配方法[J]. 电子与信息学报, 2025, 47(5): 1494–1504. doi: 10.11999/JEIT240709.

    LI Yibing, SUN Liuqing, and QI Changlong. Collaborative interference resource allocation method based on improved secretary bird algorithm[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1494–1504. doi: 10.11999/JEIT240709.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(8)  / Tables(3)

    Article Metrics

    Article views (286) PDF downloads(29) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return