Advanced Search
Turn off MathJax
Article Contents
ZHANG Xindi, CHEN Hui, ZHANG Hongyun, LIAN Feng, ZHANG Guanghua, YIN Zhipeng. Labeled Multi-Bernoulli Sensor Management Strategy Based on Twin-Delayed Deep Deterministic Policy Gradient Learning Mechanism[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260045
Citation: ZHANG Xindi, CHEN Hui, ZHANG Hongyun, LIAN Feng, ZHANG Guanghua, YIN Zhipeng. Labeled Multi-Bernoulli Sensor Management Strategy Based on Twin-Delayed Deep Deterministic Policy Gradient Learning Mechanism[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260045

Labeled Multi-Bernoulli Sensor Management Strategy Based on Twin-Delayed Deep Deterministic Policy Gradient Learning Mechanism

doi: 10.11999/JEIT260045 cstr: 32379.14.JEIT260045
Funds:  The National Natural Science Foundation of China (62163023, 61873116 ), Gansu Provincial Science and Technology Plan Project of China (25JRRA058, 25ZYJA040), 2024 Gansu Provincial Key Talent Project of China, 2023 Gansu Provincial Special Fund for Military-Civilian Integration Development of China
  • Received Date: 2026-01-13
  • Accepted Date: 2026-06-30
  • Rev Recd Date: 2026-06-30
  • Available Online: 2026-07-13
  •   Objective  Multi-target tracking requires sensor management to adapt the observation process to clutter, missed detections, target-number variations, and target-motion changes. Conventional methods typically search over a finite set of sensor actions, which increases computational cost and limits control resolution. Furthermore, reward functions constructed from multiple single-target metrics may not adequately characterize the joint multi-target posterior. To address these limitations, a continuous-action sensor management method that integrates Twin-Delayed Deep Deterministic Policy Gradient (TD3) with the Labeled Multi-Bernoulli (LMB) filter is proposed to optimize the mobile-sensor heading angle according to the multi-target belief state.  Methods  The LMB posterior, including target existence probabilities and state densities, is used to construct the belief state. At each filtering step, the mobile sensor selects a continuous heading angle that determines the sensor-target geometry, detection probability, and LMB update. Predicted target states and candidate heading actions are used to generate pseudo measurements and obtain pseudo-updated LMB densities. The Cauchy-Schwarz (CS) divergence between the predicted and pseudo-updated LMB densities is adopted to construct an information-gain reward. TD3 employs twin critics, target policy smoothing, and delayed policy updates to reduce value-estimation bias. Random control, Policy Gradient (PG)-based sensor management, CS divergence-based sensor management, and Deep Deterministic Policy Gradient (DDPG)-based sensor management are used for comparison.  Results and Discussions  DDPG-LMB and TD3-LMB produce smoother sensor trajectories than the discrete-action methods (Fig. 2). TD3-LMB achieves the highest or near-highest detection probabilities for most targets (Fig. 3) and yields larger CS divergence values during most time steps, while random control consistently produces lower values (Fig. 4). TD3-LMB also achieves the lowest overall Optimal Subpattern Assignment (OSPA) distance in the evaluated scenario, while DDPG-LMB generally outperforms the discrete baseline methods (Fig. 5). These results demonstrate that continuous heading-angle control improves observation quality and enhances the overall tracking performance of the LMB filter.  Conclusions  A TD3-based continuous-action sensor management framework for the LMB filter is presented. Candidate heading actions are evaluated using pseudo-updated LMB densities and CS divergence, directly associating action selection with the joint multi-target posterior. Simulation results demonstrate smoother sensor trajectories, higher detection probabilities for most targets, greater information gain, and lower OSPA distances in the evaluated scenario. Future work will consider higher-dimensional action spaces and cooperative multi-sensor management.
  • loading
  • [1]
    LIU Zhihao, SHANG Yuanyuan, LI Timing, et al. Robust multi-drone multi-target tracking to resolve target occlusion: A benchmark[J]. IEEE Transactions on Multimedia, 2023, 25: 1462–1476. doi: 10.1109/TMM.2023.3234822.
    [2]
    HU Juan, ZUO Lei, VARSHNEY P K, et al. Resource allocation for distributed multitarget tracking in radar networks with missing data[J]. IEEE Transactions on Signal Processing, 2024, 72: 718–734. doi: 10.1109/TSP.2024.3352915.
    [3]
    COX P B and VAN ROSSUM W L. Split-aperture phased array radar resource management for tracking tasks[J]. IEEE Transactions on Aerospace and Electronic Systems, 2025, 61(3): 6476–6486. doi: 10.1109/TAES.2025.3531843.
    [4]
    HAYKIN S. Cognitive radar: A way of the future[J]. IEEE Signal Processing Magazine, 2006, 23(1): 30–40. doi: 10.1109/MSP.2006.1593335.
    [5]
    ZHANG Peng, YAN Junkun, PU Wenqiang, et al. Multi-dimensional resource management scheme for multiple target tracking under dynamic electromagnetic environment[J]. IEEE Transactions on Signal Processing, 2024, 72: 2377–2393. doi: 10.1109/TSP.2024.3390119.
    [6]
    CHARLISH A, HOFFMANN F, DEGEN C, et al. The development from adaptive to cognitive radar resource management[J]. IEEE Aerospace and Electronic Systems Magazine, 2020, 35(6): 8–19. doi: 10.1109/MAES.2019.2957847.
    [7]
    KURNIAWATI H. Partially observable Markov decision processes and robotics[J]. Annual Review of Control, Robotics, and Autonomous Systems, 2022, 5(1): 253–277. doi: 10.1146/annurev-control-042920-092451.
    [8]
    MAHLER R. Objective functions for Bayesian control-theoretic sensor management, II: MНC-like approximation[M]. BUTENKO S, MURPHEY R, and PARDALOS P M. Recent Developments in Cooperative Control and Optimization. Boston: Springer, 2004: 273–316. doi: 10.1007/978-1-4613-0219-3_16.
    [9]
    GOSTAR A K, HOSEINNEZHAD R, BAB-HADIASHAR A, et al. OSPA-based sensor control[C]. 2015 International Conference on Control, Automation and Information Sciences (ICCAIS), Changshu, China, 2015: 214–218. doi: 10.1109/ICCAIS.2015.7338664.
    [10]
    JONES G, GARCÍA-FERNÁNDEZ Á F, and WONG P W. GOSPA-driven Gaussian Bernoulli sensor management[C]. 2023 26th International Conference on Information Fusion (FUSION), Charleston, USA, 2023: 1–8. doi: 10.23919/FUSION52260.2023.10224220.
    [11]
    GOSTAR A K, HOSEINNEZHAD R, and BAB-HADIASHAR A. Multi-Bernoulli sensor control for multi-target tracking[C]. 2013 IEEE 8th International Conference on Intelligent Sensors, Sensor Networks and Information Processing, Melbourne, Australia, 2013: 312–317. doi: 10.1109/ISSNIP.2013.6529808.
    [12]
    RISTIC B, VO B N, and CLARK D. A note on the reward function for PHD filters with sensor control[J]. IEEE Transactions on Aerospace and Electronic Systems, 2011, 47(2): 1521–1529. doi: 10.1109/TAES.2011.5751278.
    [13]
    GOSTAR A K, HOSEINNEZHAD R, RATHNAYAKE T, et al. Constrained sensor control for labeled multi-Bernoulli filter using Cauchy-Schwarz divergence[J]. IEEE Signal Processing Letters, 2017, 24(9): 1313–1317. doi: 10.1109/LSP.2017.2723924.
    [14]
    HOFFMANN F, CHARLISH A, RITCHIE M, et al. Sensor path planning using reinforcement learning[C]. 2020 IEEE 23rd International Conference on Information Fusion (FUSION), Rustenburg, South Africa, 2020: 1–8. doi: 10.23919/FUSION45008.2020.9190242.
    [15]
    SHI Yuchun, JIU Bo, YAN Junkun, et al. Data-driven simultaneous multibeam power allocation: When multiple targets tracking meets deep reinforcement learning[J]. IEEE Systems Journal, 2021, 15(1): 1264–1274. doi: 10.1109/JSYST.2020.2984774.
    [16]
    LU Ziyang and GURSOY M C. Resource allocation for multi-target radar tracking via constrained deep reinforcement learning[J]. IEEE Transactions on Cognitive Communications and Networking, 2023, 9(6): 1677–1690. doi: 10.1109/TCCN.2023.3304634.
    [17]
    OBOREH-SNAPPS O, SHE Buxin, FAHAD S, et al. Virtual synchronous generator control using twin delayed deep deterministic policy gradient method[J]. IEEE Transactions on Energy Conversion, 2024, 39(1): 214–228. doi: 10.1109/TEC.2023.3309955.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(5)

    Article Metrics

    Article views (293) PDF downloads(33) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return