← Search

Asaf Cassel

9 accepted papers

2025

Batch Ensemble for Variance Dependent Regret in Stochastic Bandits

AAAI 2025technical

Efficiently trading off exploration and exploitation is one of the key challenges in online Reinforcement Learning (RL). Most works achieve this by carefully estimating the model uncertainty and following the so-called optimistic model. Inspired by practical ensemble methods, in this work we propose…

2024

Multi-turn Reinforcement Learning with Preference Human Feedback

NeurIPS 2024poster

Reinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate remarkable abilities in various tasks. Existing methods work by emulating the human preference at the single decision (tu…

Cited by 17SourcePDFScholar
2024

Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback

ICML 2024poster

In many real-world applications, it is hard to provide a reward signal in each step of a Reinforcement Learning (RL) process and more natural to give feedback when an episode ends. To this end, we study the recently proposed model of RL with Aggregate Bandit Feedback (RL-ABF), where the agent only o…

Cited by 4SourcePDFScholar
2024

Warm-up Free Policy Optimization: Improved Regret in Linear Markov Decision Processes

NeurIPS 2024poster

Policy Optimization (PO) methods are among the most popular Reinforcement Learning (RL) algorithms in practice. Recently, Sherman et al. [2023a] proposed a PO-based algorithm with rate-optimal regret guarantees under the linear Markov Decision Process (MDP) model. However, their algorithm relies on…

Cited by 0SourcePDFScholar
2023

Efficient Rate Optimal Regret for Adversarial Contextual MDPs Using Online Function Approximation

ICML 2023poster

We present the OMG-CMDP! algorithm for regret minimization in adversarial Contextual MDPs. The algorithm operates under the minimal assumptions of realizable function class and access to online least squares and log loss regression oracles. Our algorithm is efficient (assuming efficient online regre…

Cited by 6SourcePDFScholar
2020

Bandit Linear Control

NeurIPS 2020spotlight

We consider the problem of controlling a known linear dynamical system under stochastic noise, adversarially chosen costs, and bandit feedback. Unlike the full feedback setting where the entire cost function is revealed after each decision, here only the cost incurred by the learner is observed. We…

Cited by 25SourcePDFScholar