面向博弈对抗的多智能体强化学习协同决策
DOI:
作者:
作者单位:

陆军兵种大学北京营区

作者简介:

通讯作者:

中图分类号:

TP312

基金项目:


Multi-Agent Reinforcement Learning for Cooperative Decision-Making in Adversarial Games
Author:
Affiliation:

PLA Army Services University

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    群体智能在现代军事体系中日益重要,能够显著降低作战风险、提升整体行动效能。智能体之间的协同与配合是实现群体在复杂战场环境中获得全面优势的关键。本文提出一种名为TransGMix的多智能体强化学习协同决策模型,该模型由决策模型和混合模型两部分构成: 在决策模型中,引入Transformer对智能体的局部观测进行表征映射,从而得到各个体的策略;进一步设计决策经验更新机制,利用历史决策对当前策略进行时序增强,以提升决策的稳定性与准确性。在混合模型中,结合图神经网络与多层感知器构建超网络,通过价值分解促进个体间的信息整合与群体层面的协同优化,实现从个体理性到群体协调的统一。本文在多场景仿真环境(如《星际争霸》)中对TransGMix进行了验证,取得94.5%的平均胜率,表明该方法能够有效刻画群体协同涌现与全局优化能力。

    Abstract:

    Group intelligence is becoming increasingly important in modern military systems, as it can significantly reduce operational risks and improve overall mission effectiveness. Coordination and cooperation among agents are critical for enabling a group to achieve comprehensive advantages in complex battlefield environments. In this paper, a cooperative decision-making model for multi-agent reinforcement learning, termed TransGMix, is proposed. The model consists of two main components: a decision model and a mixing model. In the decision model, a Transformer is introduced to encode the agents’ local observations, through which individual policies are generated. Furthermore, a decision experience update mechanism is designed to enhance temporal consistency by incorporating historical decision information, thereby improving the stability and accuracy of the current decision process. In the mixing model, a hypernetwork is constructed by integrating graph neural networks and multilayer perceptrons. Through value decomposition, the proposed structure facilitates information integration among agents and promotes cooperative optimization at the group level, enabling a unified transition from individual rationality to collective coordination. Experiments conducted in multi-scenario simulation environments, such as the StarCraft Multi-Agent Challenge, demonstrate that TransGMix achieves an average win rate of 94.5%, indicating that the proposed method can effectively capture emergent cooperation and global optimization capabilities in multi-agent systems.

    参考文献
    相似文献
    引证文献
引用本文

孙子文,焦生磊,孙岩. 面向博弈对抗的多智能体强化学习协同决策[J]. 科学技术与工程, , ():

复制
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-11
  • 最后修改日期:2026-05-08
  • 录用日期:2026-06-22
  • 在线发布日期:
  • 出版日期:
×
2026年会通知 | “技术经济学驱动智能经济生态构建与治理变革”——中国技术经济学会第三十三届学术年会(2026)会议通知暨征文启事(第一轮)
亟待确认版面费归属稿件,敬请作者关注