Warm-Starting Nested Rollout Policy Adaptation with Optimal Stopping

Chen Dang; Cristina Bazgan; Tristan Cazenave; Morgan Chopin; Pierre-Henri Wuillemin

doi:10.1609/aaai.v37i10.26459

Authors

Chen Dang Orange Labs, Châtillon, France Université Paris-Dauphine, PSL Research University, CNRS, UMR 7243, LAMSADE, F-75016 Paris, France
Cristina Bazgan Université Paris-Dauphine, PSL Research University, CNRS, UMR 7243, LAMSADE, F-75016 Paris, France
Tristan Cazenave Université Paris-Dauphine, PSL Research University, CNRS, UMR 7243, LAMSADE, F-75016 Paris, France
Morgan Chopin Orange Labs, Châtillon, France
Pierre-Henri Wuillemin Sorbonne Université, CNRS, UMR 7606, LIP6, F-75005 Paris, France

DOI:

https://doi.org/10.1609/aaai.v37i10.26459

Keywords:

SO: Sampling/Simulation-Based Search, PRS: Routing

Abstract

Nested Rollout Policy Adaptation (NRPA) is an approach using online learning policies in a nested structure. It has achieved a great result in a variety of difficult combinatorial optimization problems. In this paper, we propose Meta-NRPA, which combines optimal stopping theory with NRPA for warm-starting and significantly improves the performance of NRPA. We also present several exploratory techniques for NRPA which enable it to perform better exploration. We establish this for three notoriously difficult problems ranging from telecommunication, transportation and coding theory namely Minimum Congestion Shortest Path Routing, Traveling Salesman Problem with Time Windows and Snake-in-the-Box. We also improve the lower bounds of the Snake-in-the-Box problem for multiple dimensions.

Warm-Starting Nested Rollout Policy Adaptation with Optimal Stopping

Authors

DOI:

Keywords:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information

Subscription