• 中国期刊全文数据库
  • 中国学术期刊综合评价数据库
  • 中国科技论文与引文数据库
  • 中国核心期刊(遴选)数据库
徐鼎, 赵岭忠, 翟仲毅. 基于集成的多智能体深度强化学习决策优化方法J. 桂林电子科技大学学报, 2026, 46(3): 263-270. DOI: 10.16725/j.1673-808X.2023106
引用本文: 徐鼎, 赵岭忠, 翟仲毅. 基于集成的多智能体深度强化学习决策优化方法J. 桂林电子科技大学学报, 2026, 46(3): 263-270. DOI: 10.16725/j.1673-808X.2023106
Xu Ding, Zhao Lingzhong, Zhai Zhongyi. Ensemble-based decision optimization method for multi-agent reinforcement learningJ. Journal of Guilin University of Electronic Technology, 2026, 46(3): 263-270. DOI: 10.16725/j.1673-808X.2023106
Citation: Xu Ding, Zhao Lingzhong, Zhai Zhongyi. Ensemble-based decision optimization method for multi-agent reinforcement learningJ. Journal of Guilin University of Electronic Technology, 2026, 46(3): 263-270. DOI: 10.16725/j.1673-808X.2023106

基于集成的多智能体深度强化学习决策优化方法

Ensemble-based decision optimization method for multi-agent reinforcement learning

  • 摘要: 在中心训练分散执行的框架下,智能体虽然在训练期间可以获取全局状态,但是所有智能体都在同时更新自己的策略,因此会存在环境非平稳性问题,从而导致智能体在某些状态下的决策受到影响。此外,许多方法通过经验回放机制来减少样本相关性,但同时可能会导致某些状态下的样本未被抽取出来以用于训练,进而导致智能体在这些状态下决策效果不佳。以上问题导致多智能体深度强化学习中智能体的决策鲁棒性不足。针对该问题,提出了一种基于集成的多智能体深度强化学习决策优化方法。该方法通过给智能体集成多个策略网络或者动作值网络来提升智能体决策的鲁棒性,当某个状态下单个策略或动作网络失效时,智能体仍可依据其他正常工作的网络进行决策。另外,基于强化学习的特性提出了基于动作置信度权重的集成方法和基于最优动作投票的集成方法,能有效对多个网络的输出进行集成。该方法在多智能体深度强化学习领域具有广泛的适用性,既适用于基于演员−评论家的方法,又适用于基于值函数分解的方法。在3种不同的实验环境中进行了实验,对6种方法进行了对比测试。实验结果表明,将该方法应用到主流的演员−评论家方法和值函数分解方法时,能使多智能体系统的对战胜率提升50个百分点,奖励值最高提升40。

     

    Abstract: In the centralized training and decentralized execution(CTDE) framework, agents can access the global state during training, but update their policies independently. This results in environmental non-stationarity, which degrades decision-making performance in certain states. Additionally, experience replay is commonly adopted to reduce sample correlation, but it may lead to insufficient sampling of certain states, resulting in poor decision-making performance. These issues reduce the robustness of decision-making in multi-agent deep reinforcement learning. To address these issues, an ensemble-based multi-agent deep reinforcement learning method is proposed. The proposed method integrates multiple policy or action-value networks to enhance decision robustness. When one network performs poorly in a particular state, decisions can still be made based on other effective networks. Two integration strategies, namely confidence-weighted action selection and action voting, are proposed to effectively combine outputs from multiple networks. The proposed approach demonstrates strong applicability in multi-agent deep reinforcement learning and applies to both actor-critic and value-based methods. Experiments were conducted in three environments by comparing the proposed method with six representative approaches. The results show that, when applied to the mainstream actor-critic algorithm MADDPG and the value decomposition method QMIX, the proposed method increases the win rate of the multi-agent system by 50%, and improves the maximum reward by 40.

     

/

返回文章
返回