← Search

Kefan Dong

11 accepted papers

2023

Beyond NTK with Vanilla Gradient Descent: A Mean-Field Analysis of Neural Networks with Polynomial Width, Samples, and Time

NeurIPS 2023poster

Despite recent theoretical progress on the non-convex optimization of two-layer neural networks, it is still an open question whether gradient descent on neural networks without unnatural modifications can achieve better sample complexity than kernel methods. This paper provides a clean mean-field a…

Cited by 15SourcePDFScholar
2023

Model-Based Offline Reinforcement Learning with Local Misspecification

AAAI 2023technical

We present a model-based offline reinforcement learning policy performance lower bound that explicitly captures dynamics model misspecification and distribution mismatch and we propose an empirical algorithm for optimal offline policy selection. Theoretically, we prove a novel safe policy improvemen…

Cited by 4SourcePDFScholar
2021

Design of Experiments for Stochastic Contextual Linear Bandits

NeurIPS 2021poster

In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, there can be a significant engineering overhead to deploy these algorithms, especially when the dataset is collected in a…

Cited by 33SourcePDFScholar
2021

Provable Model-based Nonlinear Bandit and Reinforcement Learning: Shelve Optimism, Embrace Virtual Curvature

NeurIPS 2021poster

This paper studies model-based bandit and reinforcement learning (RL) with nonlinear function approximations. We propose to study convergence to approximate local maxima because we show that global convergence is statistically intractable even for one-layer neural net bandit with a deterministic rew…

Cited by 47SourcePDFScholar
2020

On the Expressivity of Neural Networks for Deep Reinforcement Learning

ICML 2020poster

We compare the model-free reinforcement learning with the model-based approaches through the lens of the expressive power of neural networks for policies, Q-functions, and dynamics. We show, theoretically and empirically, that even for one-dimensional continuous state space, there are many MDPs whos…

2020

Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP

ICLR 2020poster

A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. (2018) proposed a Q-learning algorithm with UCB exploration policy, and proved it has nearly optimal regret bound for finite-horizon episodic MDP. In this paper, we adapt Q-l…

Cited by 125SourceScholar