← Search

Orin Levy

8 accepted papers

2026

Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation

ICML 2026poster

We introduce OPO-CMDP, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our approach achieves a high probability regret bound of $\widetilde{O}(H^4\sqrt{T|S||A|\log(|\mathcal{F}||\mathcal{P}|)}),$ where $S…

Cited by 0SourceScholar
2025

Batch Ensemble for Variance Dependent Regret in Stochastic Bandits

AAAI 2025technical

Efficiently trading off exploration and exploitation is one of the key challenges in online Reinforcement Learning (RL). Most works achieve this by carefully estimating the model uncertainty and following the so-called optimistic model. Inspired by practical ensemble methods, in this work we propose…

2025

Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback

NeurIPS 2025spotlight

We present regret minimization algorithms for the contextual multi-armed bandit (CMAB) problem over $K$ actions in the presence of delayed feedback, a scenario where loss observations arrive with delays chosen by an adversary. As a preliminary result, assuming direct access to a finite policy clas…

Cited by 0SourceScholar
2023

Efficient Rate Optimal Regret for Adversarial Contextual MDPs Using Online Function Approximation

ICML 2023poster

We present the OMG-CMDP! algorithm for regret minimization in adversarial Contextual MDPs. The algorithm operates under the minimal assumptions of realizable function class and access to online least squares and log loss regression oracles. Our algorithm is efficient (assuming efficient online regre…

Cited by 6SourcePDFScholar