Multi-Agent Reinforcement Learning is A Sequence Modeling Problem

被引:0
|
作者
Wen, Muning [1 ,2 ]
Kuba, Jakub Grudzien [3 ]
Lin, Runji [4 ]
Zhang, Weinan [1 ]
Wen, Ying [1 ]
Wang, Jun [2 ,5 ]
Yang, Yaodong [6 ,7 ]
机构
[1] Shanghai Jiao Tong Univ, Shanghai, Peoples R China
[2] Digital Brain Lab, Berkeley, CA USA
[3] Univ Oxford, Oxford, England
[4] Chinese Acad Sci, Inst Automat, Beijing, Peoples R China
[5] UCL, London, England
[6] Beijing Inst Gen AI, Beijing, Peoples R China
[7] Peking Univ, Inst AI, Beijing, Peoples R China
基金
中国国家自然科学基金;
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Large sequence models (SM) such as GPT series and BERT have displayed outstanding performance and generalization capabilities in natural language process, vision and recently reinforcement learning. A natural follow-up question is how to abstract multi-agent decision making also as an sequence modeling problem and benefit from the prosperous development of the SMs. In this paper, we introduce a novel architecture named Multi-Agent Transformer (MAT) that effectively casts co-operative multi-agent reinforcement learning (MARL) into SM problems wherein the objective is to map agents' observation sequences to agents' optimal action sequences. Our goal is to build the bridge between MARL and SMs so that the modeling power of modern sequence models can be unleashed for MARL. Central to our MAT is an encoder-decoder architecture which leverages the multi-agent advantage decomposition theorem to transform the joint policy search problem into a sequential decision making process; this renders only linear time complexity for multi-agent problems and, most importantly, endows MAT with monotonic performance improvement guarantee. Unlike prior arts such as Decision Transformer fit only pre-collected offline data, MAT is trained by online trial and error from the environment in an on-policy fashion. To validate MAT, we conduct extensive experiments on StarCraftII, Multi-Agent MuJoCo, Dexterous Hands Manipulation, and Google Research Football benchmarks. Results demonstrate that MAT achieves superior performance and data efficiency compared to strong baselines including MAPPO and HAPPO. Furthermore, we demonstrate that MAT is an excellent few-short learner on unseen tasks regardless of changes in the number of agents. See our project page at https://sites.google.com/view/multi-agent-transformer((1)).
引用
收藏
页数:13
相关论文
共 50 条
  • [31] Consensus Learning for Cooperative Multi-Agent Reinforcement Learning
    Xu, Zhiwei
    Zhang, Bin
    Li, Dapeng
    Zhang, Zeren
    Zhou, Guangchong
    Chen, Hao
    Fan, Guoliang
    THIRTY-SEVENTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 37 NO 10, 2023, : 11726 - 11734
  • [32] Concept Learning for Interpretable Multi-Agent Reinforcement Learning
    Zabounidis, Renos
    Campbell, Joseph
    Stepputtis, Simon
    Hughes, Dana
    Sycara, Katia
    CONFERENCE ON ROBOT LEARNING, VOL 205, 2022, 205 : 1828 - 1837
  • [33] Learning structured communication for multi-agent reinforcement learning
    Sheng, Junjie
    Wang, Xiangfeng
    Jin, Bo
    Yan, Junchi
    Li, Wenhao
    Chang, Tsung-Hui
    Wang, Jun
    Zha, Hongyuan
    AUTONOMOUS AGENTS AND MULTI-AGENT SYSTEMS, 2022, 36 (02)
  • [34] Learning structured communication for multi-agent reinforcement learning
    Junjie Sheng
    Xiangfeng Wang
    Bo Jin
    Junchi Yan
    Wenhao Li
    Tsung-Hui Chang
    Jun Wang
    Hongyuan Zha
    Autonomous Agents and Multi-Agent Systems, 2022, 36
  • [35] Generalized learning automata for multi-agent reinforcement learning
    De Hauwere, Yann-Michael
    Vrancx, Peter
    Nowe, Ann
    AI COMMUNICATIONS, 2010, 23 (04) : 311 - 324
  • [36] Network-aware Multi-agent Reinforcement Learning for the Vehicle Navigation Problem
    Arasteh, Fazel
    SheikhGarGar, Soroush
    Papagelis, Manos
    30TH ACM SIGSPATIAL INTERNATIONAL CONFERENCE ON ADVANCES IN GEOGRAPHIC INFORMATION SYSTEMS, ACM SIGSPATIAL GIS 2022, 2022, : 504 - 507
  • [37] A Multi-Agent Reinforcement Learning Approach to the Dynamic Job Shop Scheduling Problem
    Inal, Ali Firat
    Sel, Cagri
    Aktepe, Adnan
    Turker, Ahmet Kursad
    Ersoz, Suleyman
    SUSTAINABILITY, 2023, 15 (10)
  • [38] Multi-Agent Reinforcement Learning with Shared Policy for Cloud Quota Management Problem
    Cheng, Tong
    Dong, Hang
    Wang, Lu
    Qiao, Bo
    Qin, Si
    Lin, Qingwei
    Zhang, Dongmei
    Rajmohan, Saravan
    Moscibroda, Thomas
    COMPANION OF THE WORLD WIDE WEB CONFERENCE, WWW 2023, 2023, : 391 - 395
  • [39] Application of Multi-agent Reinforcement Learning to the Dynamic Scheduling Problem in Manufacturing Systems
    Heik, David
    Bahrpeyma, Fouad
    Reichelt, Dirk
    MACHINE LEARNING, OPTIMIZATION, AND DATA SCIENCE, LOD 2023, PT II, 2024, 14506 : 237 - 254
  • [40] Multi-agent reinforcement learning for character control
    Li, Cheng
    Fussell, Levi
    Komura, Taku
    VISUAL COMPUTER, 2021, 37 (12): : 3115 - 3123