상세 보기
D-QMIX: multi-step sequential forward dynamics modeling with global state and self-attention for sample-efficient multi-agent reinforcement learning
- Kim, Jung In;
- Kim, Seoung Bum
WEB OF SCIENCE
2SCOPUS
2초록
Improving sample efficiency under limited interactions is a critical challenge in multi-agent reinforcement learning (MARL). Although existing studies such as QMIX have successfully modeled agent cooperation through centralized training with decentralized execution, they require large amounts of data to perform effectively, limiting real-world applicability. To overcome this challenge, we propose D-QMIX, which uses multi-step sequential forward dynamics modeling (MSFDM) as an auxiliary task for QMIX. Unlike conventional approaches that predict only the immediate next observation, D-QMIX predicts future observations further ahead by using sequential inputs through an MSFDM, which consists of three key components: (1) forward dynamics modeling to learn environment dynamics, (2) a self-attention mechanism to focus on key inter-agent correlations, and (3) the use of global state information to mitigate prediction inaccuracies arising from local observations. Experimental results on the StarCraft II micromanagement benchmark (SMAC) and the multi-agent particle environment (MPE) demonstrate the superior sample efficiency of D-QMIX. Specifically, D-QMIX outperforms existing methods in eight out of 12 scenarios in SMAC and achieves higher performance in two out of four scenarios in MPE. These results validate the effectiveness of integrating an MSFDM with QMIX to enhance sample efficiency across various MARL environments. The code is available at https://github. com/junginkim23/D-QMIX.
키워드
- 제목
- D-QMIX: multi-step sequential forward dynamics modeling with global state and self-attention for sample-efficient multi-agent reinforcement learning
- 저자
- Kim, Jung In; Kim, Seoung Bum
- 발행일
- 2026-03
- 유형
- Article
- 권
- 729