TY - JOUR AU - Ramanujan, Raghuram AU - Selman, Bart PY - 2011/03/22 Y2 - 2024/03/29 TI - Trade-Offs in Sampling-Based Adversarial Planning JF - Proceedings of the International Conference on Automated Planning and Scheduling JA - ICAPS VL - 21 IS - 1 SE - Full Technical Papers DO - 10.1609/icaps.v21i1.13472 UR - https://ojs.aaai.org/index.php/ICAPS/article/view/13472 SP - 202-209 AB - <p> The Upper Confidence bounds for Trees (UCT) algorithm has in recent years captured the attention of the planning and game-playing community due to its notable success in the game of Go. However, attempts to reproduce similar levels of performance in domains that are the forte of Minimax-style algorithms have been largely unsuccessful, making any comparative studies of the two hard. In this paper, we study UCT in the game of Mancala, which to our knowledge is the first domain where both search algorithms perform quite well with minimal enhancement. We focus on the three key components of the UCT algorithm in its purest form - targeted node expansion, state value estimation via playouts and averaging backups - and look at their contributions to the overall performance of the algorithm. We study the trade-offs involved in using alternate ways to perform these steps. Finally, we demonstrate a novel hybrid approach to enhancing UCT, that exploits its superior decision accuracy in regions of the search space with few terminal nodes. </p> ER -