基于多智能体博弈信息对抗学习的行人-车辆交互行为建模仿真
DOI:
作者:
作者单位:

重庆理工大学

作者简介:

通讯作者:

中图分类号:

U491.2+25

基金项目:

重庆市自然科学基金面上项目(CSTB2025NSCQ-GPX0222)


Multi Agent Game Information Adversarial Inverse Reinforcement Learning for Pedestrian-Vehicle Interaction Behavior Modeling
Author:
Affiliation:

Chongqing University of Technology

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对复杂城市交通场景中行人-车辆动态交互难以高保真仿真的问题,本文提出一种基于博弈论与信息论的多智能体对抗逆强化学习(multi-agent game information adversarial inverse reinforcement learning, MA-GI-AIRL)模型。首先,采用逻辑随机最佳响应均衡对有限理性下的次优决策行为进行建模,并基于碰撞时间(time-to-collision, TTC)、后侵入时间(post-encroachment time, PET)、速度和加速度等运动学指标,分别表征交互过程中的安全性、通行效率与运动舒适性,进而构建结构化博弈效用函数;其次,引入互信息最大化机制以解耦异质行为偏好;最后,基于无人机采集的真实交通数据,学习并恢复可解释的奖励函数与行为策略。研究结果表明,MA-GI-AIRL模型在轨迹拟合精度(行人平均位移误差为0.757 m)和运动学指标合理性方面均优于基准模型,为自动驾驶系统研发提供高保真的人车交互仿真环境。

    Abstract:

    High-fidelity simulation of pedestrian-vehicle dynamic interaction in complex urban traffic scenarios is identified as a critical unsolved technical challenge in the autonomous driving field. To address this challenge, a multi-agent game information adversarial inverse reinforcement learning (MA-GI-AIRL) model was proposed in this work. First, the logit stochastic best-response equilibrium was adopted to model suboptimal decision-making behaviors under bounded rationality, and a structured game utility function was constructed by using kinematic indicators such as time-to-collision (TTC), post-encroachment time (PET), velocity, and acceleration to characterize interaction safety, traffic efficiency, and motion comfort, respectively. Second, a mutual information maximization mechanism was introduced to decouple heterogeneous behavioral preferences. Finally, real-world traffic data collected by unmanned aerial vehicles was utilized for model training. An interpretable reward function and corresponding behavioral policies were learned and recovered. The research results show that the proposed MA-GI-AIRL model outperforms the baseline models in terms of both trajectory fitting accuracy and the rationality of kinematic indicators. The average displacement error (ADE) of the model is 0.757 m for pedestrian trajectories. A high-fidelity pedestrian-vehicle interaction simulation environment is provided for development and evaluation of autonomous driving systems by the validated model.

    参考文献
    相似文献
    引证文献
引用本文

王云卿,雷文浩,李文礼. 基于多智能体博弈信息对抗学习的行人-车辆交互行为建模仿真[J]. 科学技术与工程, , ():

复制
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-03-05
  • 最后修改日期:2026-06-05
  • 录用日期:2026-07-27
  • 在线发布日期:
  • 出版日期:
×
2026年会通知 | “技术经济学驱动智能经济生态构建与治理变革”——中国技术经济学会第三十三届学术年会(2026)会议通知暨征文启事(第一轮)
亟待确认版面费归属稿件,敬请作者关注