← Search

Botao Hao

17 accepted papers

2023

Leveraging Demonstrations to Improve Online Learning: Quality Matters

ICML 2023poster

We investigate the extent to which offline demonstration data can improve online learning. It is natural to expect some improvement, but *the question is how, and by how much?* We show that the degree of improvement must depend on the *quality* of the demonstration data. To generate portable insight…

Cited by 9SourcePDFScholar
2022

Confident Least Square Value Iteration with Local Access to a Simulator

AISTATS 2022poster

Learning with simulators is ubiquitous in mod-ern reinforcement learning (RL). The simulatorcan either correspond to a simplified version ofthe real environment (such as a physics simulation of a robot arm) or to the environment itself (such as in games like Atari and Go). Among algorithms that are…

Cited by 9SourcePDFScholar
2022

Interacting Contour Stochastic Gradient Langevin Dynamics

ICLR 2022poster

We propose an interacting contour stochastic gradient Langevin dynamics (ICSGLD) sampler, an embarrassingly parallel multiple-chain contour stochastic gradient Langevin dynamics (CSGLD) sampler with efficient interactions. We show that ICSGLD can be theoretically more efficient than a single-chain C…

2022

The Neural Testbed: Evaluating Joint Predictions

NeurIPS 2022accept

Predictive distributions quantify uncertainties ignored by point estimates. This paper introduces The Neural Testbed: an open source benchmark for controlled and principled evaluation of agents that generate such predictions. Crucially, the testbed assesses agents not only on the quality of their ma…

2021

Adaptive Approximate Policy Iteration

AISTATS 2021poster

Model-free reinforcement learning algorithms combined with value function approximation have recently achieved impressive performance in a variety of application domains. However, the theoretical understanding of such algorithms is limited, and existing results are largely focused on episodic or dis…

Cited by 15SourcePDFScholar
2021

Bandit Phase Retrieval

NeurIPS 2021poster

We study a bandit version of phase retrieval where the learner chooses actions $(A_t)_{t=1}^n$ in the $d$-dimensional unit ball and the expected reward is $\langle A_t, \theta_\star \rangle^2$ with $\theta_\star \in \mathbb R^d$ an unknown parameter vector. We prove an upper bound on the minimax cum…

Cited by 12SourcePDFScholar
2021

Bootstrapping Fitted Q-Evaluation for Off-Policy Inference

ICML 2021spotlight

Bootstrapping provides a flexible and effective approach for assessing the quality of batch reinforcement learning, yet its theoretical properties are poorly understood. In this paper, we study the use of bootstrapping in off-policy evaluation (OPE), and in particular, we focus on the fitted Q-evalu…

Cited by 52SourcePDFScholar
2021

Sparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient

ICML 2021spotlight

This paper provides a statistical analysis of high-dimensional batch reinforcement learning (RL) using sparse linear function approximation. When there is a large number of candidate features, our result sheds light on the fact that sparsity-aware methods can make batch RL more sample efficient. We…

Cited by 39SourcePDFScholar