← Search

Gellért Weisz

4 accepted papers

2024

Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^\pi$-Realizability and Concentrability

NeurIPS 2024poster

We consider offline reinforcement learning (RL) in $H$-horizon Markov decision processes (MDPs) under the linear $q^\pi$-realizability assumption, where the action-value function of every policy is linear with respect to a given $d$-dimensional feature function. The hope in this setting is that lear…

Cited by 0SourcePDFScholar
2023

Online RL in Linearly $q^\pi$-Realizable MDPs Is as Easy as in Linear MDPs If You Learn What to Ignore

NeurIPS 2023oral

We consider online reinforcement learning (RL) in episodic Markov decision processes (MDPs) under the linear $q^\pi$-realizability assumption, where it is assumed that the action-values of all policies can be expressed as linear functions of state-action features. This class is known to be more gen…

Cited by 8SourcePDFScholar
2023

Optimistic Natural Policy Gradient: a Simple Efficient Policy Optimization Framework for Online RL

NeurIPS 2023spotlight

While policy optimization algorithms have played an important role in recent empirical success of Reinforcement Learning (RL), the existing theoretical understanding of policy optimization remains rather limited---they are either restricted to tabular MDPs or suffer from highly suboptimal sample com…

Cited by 9SourcePDFScholar
2022

Confident Approximate Policy Iteration for Efficient Local Planning in $q^\pi$-realizable MDPs

NeurIPS 2022accept

We consider approximate dynamic programming in $\gamma$-discounted Markov decision processes and apply it to approximate planning with linear value-function approximation. Our first contribution is a new variant of Approximate Policy Iteration (API), called Confident Approximate Policy Iteration (CA…

Cited by 13SourcePDFScholar