• 中国期刊全文数据库
  • 中国学术期刊综合评价数据库
  • 中国科技论文与引文数据库
  • 中国核心期刊(遴选)数据库
张琼, 赵岭忠, 翟仲毅. 基于QMIX的多智能体自适应探索方法J. 桂林电子科技大学学报, 2026, 46(3): 271-277. DOI: 10.16725/j.1673-808X.202356
引用本文: 张琼, 赵岭忠, 翟仲毅. 基于QMIX的多智能体自适应探索方法J. 桂林电子科技大学学报, 2026, 46(3): 271-277. DOI: 10.16725/j.1673-808X.202356
Zhang Qiong, Zhao Lingzhong, Zhai Zhongyi. Adaptive exploration for multi-agent systems based on QMIXJ. Journal of Guilin University of Electronic Technology, 2026, 46(3): 271-277. DOI: 10.16725/j.1673-808X.202356
Citation: Zhang Qiong, Zhao Lingzhong, Zhai Zhongyi. Adaptive exploration for multi-agent systems based on QMIXJ. Journal of Guilin University of Electronic Technology, 2026, 46(3): 271-277. DOI: 10.16725/j.1673-808X.202356

基于QMIX的多智能体自适应探索方法

Adaptive exploration for multi-agent systems based on QMIX

  • 摘要: 针对目前多智能体强化学习方法中智能体通常采用贪心策略易导致对环境探索不充分而陷入次优解,以及同一环境中所有智能体探索率相同而不足以适应复杂任务中的环境动态变化和其他智能体行为的问题,提出了一种结合深度强化学习方法的多智能体自适应探索方法。该方法基于QMIX算法,在2个时间层面上采用不同的方法调整智能体对环境的探索:在每个时间步计算智能体的策略熵,依据策略熵的大小按概率分布或状态动作值选择动作;定义智能体在一段时间内的“表现”以及对环境的“探索程度”,智能体参考整个团队在这段时间内的表现以及对环境的探索程度调整自己的探索率。在《星际争霸Ⅱ》场景中的3s5z与5m_vs_6m任务上进行了实验。实验结果表明,相较于基线方法,该方法在2种任务中都获得了更高的平均胜率,显著提高了多智能体系统的学习性能。

     

    Abstract: To address insufficient exploration and premature convergence to suboptimal solutions in existing multi-agent reinforcement learning (MARL) methods, an adaptive exploration method based on deep reinforcement learning is proposed. Furthermore, using the same exploration rate for all agents limits their adaptability to dynamic environments and inter-agent interactions in complex tasks. Based on the QMIX framework, the proposed method adjusts agent exploration at two temporal scales. At each time step, the policy entropy is computed, and actions are selected according to a probability distribution derived from the state-action values. At a higher temporal scale, performance and exploration indicators are defined over a time window, and exploration rates are adaptively adjusted according to global team performance and exploration status. Experiments conducted on the 3s5z and 5m_vs_6m tasks in the StarCraft Ⅱ environment show that the proposed method consistently outperforms baseline methods and achieves higher average win rates in both tasks, significantly improving the learning performance of the multi-agent system.

     

/

返回文章
返回