Thompson Sampling for Stochastic Bandits with Graph Feedback

Aristide Tossou; Christos Dimitrakakis; Devdatt Dubhashi

doi:10.1609/aaai.v31i1.10897

Authors

Aristide Tossou Chalmers University of Technology
Christos Dimitrakakis University of Lille, and Chalmers University of Technology
Devdatt Dubhashi Chalmers University of Technology

DOI:

https://doi.org/10.1609/aaai.v31i1.10897

Keywords:

Thompson sampling, Stochastic multi-armed bandit, graphical learning, online learning

Abstract

We present a simple set of algorithms based on Thompson Sampling for stochastic bandit problems with graph feedback. Thompson Sampling is generally applicable, without the need to construct complicated upper confidence bounds. As we show in this paper, it has excellent performance in problems with graph feedback, even when the graph structure itself is unknown and/or changing. We provide theoretical guarantees on the Bayesian regret of the algorithm, as well as extensive experi- mental results on real and simulated networks. More specifically, we tested our algorithms on power law, planted partitions and Erdo's–Rényi graphs, as well as on graphs derived from Facebook and Flixster data and show that they clearly outperform related methods that employ upper confidence bounds.

Thompson Sampling for Stochastic Bandits with Graph Feedback

Authors

DOI:

Keywords:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information