D-QMIX: multi-step sequential forward dynamics modeling with global state and self-attention for sample-efficient multi-agent reinforcement learning

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

2

초록

Improving sample efficiency under limited interactions is a critical challenge in multi-agent reinforcement learning (MARL). Although existing studies such as QMIX have successfully modeled agent cooperation through centralized training with decentralized execution, they require large amounts of data to perform effectively, limiting real-world applicability. To overcome this challenge, we propose D-QMIX, which uses multi-step sequential forward dynamics modeling (MSFDM) as an auxiliary task for QMIX. Unlike conventional approaches that predict only the immediate next observation, D-QMIX predicts future observations further ahead by using sequential inputs through an MSFDM, which consists of three key components: (1) forward dynamics modeling to learn environment dynamics, (2) a self-attention mechanism to focus on key inter-agent correlations, and (3) the use of global state information to mitigate prediction inaccuracies arising from local observations. Experimental results on the StarCraft II micromanagement benchmark (SMAC) and the multi-agent particle environment (MPE) demonstrate the superior sample efficiency of D-QMIX. Specifically, D-QMIX outperforms existing methods in eight out of 12 scenarios in SMAC and achieves higher performance in two out of four scenarios in MPE. These results validate the effectiveness of integrating an MSFDM with QMIX to enhance sample efficiency across various MARL environments. The code is available at https://github. com/junginkim23/D-QMIX.

키워드

Forward dynamics model; Multi-agent reinforcement learning; Representation learning; Sample efficiency; StarCraft II
제목
D-QMIX: multi-step sequential forward dynamics modeling with global state and self-attention for sample-efficient multi-agent reinforcement learning
저자
Kim, Jung In; Kim, Seoung Bum
DOI
10.1016/j.ins.2025.122867
발행일
2026-03
유형
Article
저널명
Information Sciences
권
729