基于柔性演员-评论家的车联网联邦学习客户端选择优化方法

Optimization Method for Federated Learning Client Selection in Vehicle-to-Everything Based on Soft Actor-Critic

  • 摘要: 在通信与计算资源受限的车载边缘计算环境下,针对车联网联邦学习(Federated Learning, FL)客户端选择不当,导致其聚合后全局模型性能差、训练时延高的问题,提出了一种基于柔性演员-评论家(Soft Actor-Critic, SAC)的客户端选择优化方法。首先,该方法建立了融合车辆信道状态、车辆位置与本地计算资源的系统模型,以刻画车载边缘计算环境下影响联邦学习性能的关键因素;其次,将车辆选择过程建模为马尔可夫决策过程,在连续动作空间下实现决策控制,确立了优化车辆选择决策以提升联邦学习性能的长期目标;最后,设计了一种均衡通信、计算成本与联邦损失的奖励函数,并提出对应的基于SAC的车联网联邦学习客户端选择算法。算法内部引入最大熵强化学习框架和双Critic网络结构,并通过Actor网络输出车辆参与决策,根据环境奖励引导策略实现模型性能与时延的协同优化。数值仿真结果表明,该方法充分考虑了车联网环境中车辆移动、信道条件动态变化以及计算能力异构等实际特性,能够有效选择高质量客户端参与联邦学习。对应算法相较传统FL算法、FedCS算法和DDPG算法收敛速度更快,模型精度显著提高,平均时延减少约11.56%。同时本文算法保持了较好的客户端选择多样性,增强模型在车联网场景下的泛化能力。

     

    Abstract: In vehicular edge-computing environments with limited communication and computation resources, improper client selection in federated learning (FL) for vehicle-to-everything (V2X) often causes poor performance and high training delays in the aggregated global model. In this study, we propose an optimization method for client selection based on soft actor-critic (SAC). First, this method establishes a system model that integrates the vehicle channel state, vehicle location, and local computing resources to characterize the key factors affecting the FL performance in the vehicular edge-computing environment. Next, the vehicle-selection process is modeled as a Markov decision process to achieve decision control in the continuous action space. Optimization of vehicle-selection decisions is also established to enhance the long-term goal of high FL performance. Finally, a reward function is designed to balance communication, computing cost, and FL loss, and a corresponding client-selection algorithm is presented for V2X FL based on SAC. This algorithm incorporates a maximum-entropy reinforcement learning framework and dual-critic network architecture and outputs vehicle-participation decisions through the actor network. Furthermore, the proposed algorithm realizes collaborative optimization of the model performance and delay by leveraging environmental reward guidance strategies. The numerical simulation results demonstrate that the proposed method fully considers the practical characteristics of vehicle mobility, dynamic channel conditions, and heterogeneous computational capabilities in V2X environments. The proposed method can effectively select high-quality clients to participate in FL. Compared with the traditional FL, FedCS, and DDPG algorithms, the proposed algorithm achieves faster convergence, significantly improves the model accuracy, and reduces the average training delay by approximately 11.56%. Meanwhile, the proposed algorithm maintains client-selection diversity and enhances the generalization capability of FL models in V2X scenarios.

     

/

返回文章
返回