Value-Decomposition Multi-Agent Actor-Critics

Jianyu Su; Stephen Adams; Peter Beling

doi:10.1609/aaai.v35i13.17353

Authors

Jianyu Su University of Virginia
Stephen Adams University of Virginia
Peter Beling University of Virginia

DOI:

https://doi.org/10.1609/aaai.v35i13.17353

Keywords:

Multiagent Learning, Reinforcement Learning

Abstract

The exploitation of extra state information has been an active research area in multi-agent reinforcement learning (MARL). QMIX represents the joint action-value using a non-negative function approximator and achieves the best performance on the StarCraft II micromanagement testbed, a common MARL benchmark. However, our experiments demonstrate that, in some cases, QMIX performs sub-optimally with the A2C framework, a training paradigm that promotes algorithm training efficiency. To obtain a reasonable trade-off between training efficiency and algorithm performance, we extend value-decomposition to actor-critic methods that are compatible with A2C and propose a novel actor-critic framework, value-decomposition actor-critic (VDAC). We evaluate VDAC on the StarCraft II micromanagement task and demonstrate that the proposed framework improves median performance over other actor-critic methods. Furthermore, we use a set of ablation experiments to identify the key factors that contribute to the performance of VDAC.

Value-Decomposition Multi-Agent Actor-Critics

Authors

DOI:

Keywords:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information