← Search

reda ouhamma

6 accepted papers

2025

Efficient Preference-Based Reinforcement Learning: Randomized Exploration meets Experimental Design

NeurIPS 2025poster

We study reinforcement learning from human feedback in general Markov decision processes, where agents learn from trajectory-level preference comparisons. A central challenge in this setting is to design algorithms that select informative preference queries to identify the underlying reward while en…

Cited by 0SourceScholar
2023

Bilinear Exponential Family of MDPs: Frequentist Regret Bound with Tractable Exploration & Planning

AAAI 2023technical

We study the problem of episodic reinforcement learning in continuous state-action spaces with unknown rewards and transitions. Specifically, we consider the setting where the rewards and transitions are modeled using parametric bilinear exponential families. We propose an algorithm, that a) uses pe…

Cited by 15SourcePDFScholar
2021

Learning Value Functions in Deep Policy Gradients using Residual Variance

ICLR 2021poster

Policy gradient algorithms have proven to be successful in diverse decision making and control tasks. However, these methods suffer from high sample complexity and instability issues. In this paper, we address these challenges by providing a different approach for training the critic in the actor-cr…

Cited by 24SourcePDFScholar
2021

Online Sign Identification: Minimization of the Number of Errors in Thresholding Bandits

NeurIPS 2021spotlight

In the fixed budget thresholding bandit problem, an algorithm sequentially allocates a budgeted number of samples to different distributions. It then predicts whether the mean of each distribution is larger or lower than a given threshold. We introduce a large family of algorithms (containing most e…

Cited by 7SourcePDFScholar
2021

Stochastic Online Linear Regression: the Forward Algorithm to Replace Ridge

NeurIPS 2021poster

We consider the problem of online linear regression in the stochastic setting. We derive high probability regret bounds for online $\textit{ridge}$ regression and the $\textit{forward}$ algorithm. This enables us to compare online regression algorithms more accurately and eliminate assumptions of bo…

Cited by 15SourcePDFScholar