混合强化学习框架下的水下传感器网络节点动态选择方法
Hybrid Deep Reinforcement Learning for Dynamic Node Selection in Underwater Sensor Networks
-
摘要: 水下传感器网络执行目标跟踪任务时,如何在保证跟踪精度的前提下动态选择最优传感器节点以降低网络资源消耗,是该领域亟待突破的关键问题。强化学习技术为水下传感器网络的节点选择问题,提供了一种数据驱动、自学习的智能决策方案。然而,传统单一的在线学习方法因真实数据稀缺、环境适应性不足等问题,难以同时保障网络的训练效率与策略鲁棒性。针对这一挑战,本文提出一种混合强化学习框架下的水下传感器网络节点动态选择方法,该方法通过离线虚拟训练与在线真实交互相结合的混合决策机制,实现数据增强与环境自适应。首先,基于目标运动模型与传感器观测特性,推导了水下传感器网络目标跟踪的后验克拉美罗界,并以此为约束建立了以最小化激活节点数目为目标的优化模型;其次,构建了混合深度Q网络(Hybrid Deep Q-Network, Hybrid-DQN)以求解传感器节点选择模型,利用目标运动模型与测量噪声等先验知识,在真实测量间隙生成虚拟状态与观测数据,有效扩充训练样本,从而实现训练效率与策略适应性的平衡;最后,实验结果表明,所提方法在跟踪精度方面均优于单一在线强化学习及传统启发式优化算法,其收敛速度显著加快,且决策时间大幅缩短,能以较低计算开销有效满足水下传感器动态网络的实时节点选择需求。Abstract: A key challenge when performing target tracking tasks using underwater sensor networks is the dynamic selection of optimal sensor nodes to reduce resource consumption while maintaining tracking accuracy. Reinforcement learning offers a data-driven, self-learning, and intelligent decision-making solution for this node selection problem. However, traditional online learning methods face difficulties in simultaneously ensuring training efficiency and policy robustness owing to issues such as a scarcity of real data and insufficient environmental adaptability. To address this challenge, this study proposes a dynamic node selection method for underwater sensor networks using a hybrid reinforcement learning framework. The method achieves data augmentation and environmental adaptation by combining offline virtual training with online real interaction in a hybrid decision-making mechanism. First, the posterior Cramér-Rao bound for target tracking is derived based on the target motion model and sensor observation characteristics, and an optimization model aimed at minimizing the number of active nodes is established under this constraint. Subsequently, a Hybrid Deep Q-Network is constructed to solve the sensor node selection problem. Virtual states and observation data are generated during real measurement intervals by leveraging prior knowledge such as the target motion model and measurement noise, thereby effectively expanding the training sample set and balancing training efficiency with policy adaptability. Experimental results show that the proposed method outperforms both pure online reinforcement learning and traditional heuristic optimization algorithms in tracking accuracy, with a significantly accelerated convergence speed and substantially reduced decision time. This method can effectively meet the real-time node selection requirements of dynamic underwater sensor networks with a low computational overhead.
下载: