2022
Search-based Reinforcement Learning through Bandit Linear Optimization
IJCAI 2022poster
The development of AlphaZero was a breakthrough in search-based reinforcement learning, by employing a given world model in a Monte-Carlo tree search (MCTS) algorithm to incrementally learn both an action policy and a value estimation. When extending this paradigm to the setting of simultaneous move…