← Search

Ioannis Antonoglou

7 accepted papers

2022

Planning in Stochastic Environments with a Learned Model

ICLR 2022spotlight

Model-based reinforcement learning has proven highly successful. However, learning a model in isolation from its use during planning is problematic in complex environments. To date, the most effective techniques have instead combined value-equivalent model learning with powerful tree-search methods.…

Cited by 85SourcePDFScholar
2021

Learning and Planning in Complex Action Spaces

ICML 2021spotlight

Many important real-world problems have action spaces that are high-dimensional, continuous or both, making full enumeration of all possible actions infeasible. Instead, only small subsets of actions can be sampled for the purpose of policy evaluation and improvement. In this paper, we propose a gen…

Cited by 108SourcePDFScholar
2021

Machine Translation Decoding beyond Beam Search

EMNLP 2021main

Beam search is the go-to method for decoding auto-regressive machine translation models. While it yields consistent improvements in terms of BLEU, it is only concerned with finding outputs with high model likelihood, and is thus agnostic to whatever end metric or score practitioners care about. Our…

Cited by 73SourcePDFScholar
2021

Online and Offline Reinforcement Learning by Planning with a Learned Model

NeurIPS 2021spotlight

Learning efficiently from small amounts of data has long been the focus of model-based reinforcement learning, both for the online case when interacting with the environment, and the offline case when learning from a fixed dataset. However, to date no single unified algorithm could demonstrate state…

Cited by 138SourcePDFScholar
2021

Vector Quantized Models for Planning

ICML 2021spotlight

Recent developments in the field of model-based RL have proven successful in a range of environments, especially ones where planning is essential. However, such successes have been limited to deterministic fully-observed environments. We present a new approach that handles stochastic and partially-o…

Cited by 62SourcePDFScholar
2020

Monte-Carlo Tree Search as Regularized Policy Optimization

ICML 2020poster

The combination of Monte-Carlo tree search (MCTS) with deep reinforcement learning has led to groundbreaking results in artificial intelligence. However, AlphaZero, the current state-of-the-art MCTS algorithm still relies on handcrafted heuristics that are only partially understood. In this paper, w…

Cited by 97SourcePDFScholar
2018

Learning to search with MCTSnets

ICML 2018oral

Planning problems are among the most important and well-studied problems in artificial intelligence. They are most typically solved by tree search algorithms that simulate ahead into the future, evaluate future states, and back-up those evaluations to the root of a search tree. Among these algorithm…

Cited by 107SourcePDFScholar