Counterfactual Multi-Agent Policy Gradients

Jakob Foerster; Gregory Farquhar; Triantafyllos Afouras; Nantas Nardelli; Shimon Whiteson

doi:10.1609/aaai.v32i1.11794

Authors

Jakob Foerster University of Oxford
Gregory Farquhar University of Oxford
Triantafyllos Afouras University of Oxford
Nantas Nardelli University of Oxford
Shimon Whiteson University of Oxford

DOI:

https://doi.org/10.1609/aaai.v32i1.11794

Keywords:

deep reinforcement learning, multi-agent learning, actorcritic

Abstract

Many real-world problems, such as network packet routing and the coordination of autonomous vehicles, are naturally modelled as cooperative multi-agent systems. There is a great need for new reinforcement learning methods that can efficiently learn decentralised policies for such systems. To this end, we propose a new multi-agent actor-critic method called counterfactual multi-agent (COMA) policy gradients. COMA uses a centralised critic to estimate the Q-function and decentralised actors to optimise the agents' policies. In addition, to address the challenges of multi-agent credit assignment, it uses a counterfactual baseline that marginalises out a single agent's action, while keeping the other agents' actions fixed. COMA also uses a critic representation that allows the counterfactual baseline to be computed efficiently in a single forward pass. We evaluate COMA in the testbed of StarCraft unit micromanagement, using a decentralised variant with significant partial observability. COMA significantly improves average performance over other multi-agent actor-critic methods in this setting, and the best performing agents are competitive with state-of-the-art centralised controllers that get access to the full state.

Counterfactual Multi-Agent Policy Gradients

Authors

DOI:

Keywords:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information