← Search

Takashi Onishi

2 accepted papers

2022

Dropout Q-Functions for Doubly Efficient Reinforcement Learning

ICLR 2022poster

Randomized ensembled double Q-learning (REDQ) (Chen et al., 2021b) has recently achieved state-of-the-art sample efficiency on continuous-action reinforcement learning benchmarks. This superior sample efficiency is made possible by using a large Q-function ensemble. However, REDQ is much less comput…

2019

Learning Robust Options by Conditional Value at Risk Optimization

NeurIPS 2019poster

Options are generally learned by using an inaccurate environment model (or simulator), which contains uncertain model parameters. While there are several methods to learn options that are robust against the uncertainty of model parameters, these methods only consider either the worst case or the av…