← Search

Assaf Hallak

9 accepted papers

2025

RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression

CVPR 2025poster

Video encoders optimize compression for human perception by minimizing reconstruction error under bit-rate constraints. In many modern applications such as autonomous driving, an overwhelming majority of videos serve as input for AI systems performing tasks like object recognition or segmentation, r…

Cited by 0SourcePDFScholar
2023

Planning and Learning with Adaptive Lookahead

AAAI 2023technical

Some of the most powerful reinforcement learning frameworks use planning for action selection. Interestingly, their planning horizon is either fixed or determined arbitrarily by the state visitation history. Here, we expand beyond the naive fixed horizon and propose a theoretically justified strateg…

Cited by 9SourcePDFScholar
2022

On Covariate Shift of Latent Confounders in Imitation and Reinforcement Learning

ICLR 2022poster

We consider the problem of using expert data with unobserved confounders for imitation and reinforcement learning. We begin by defining the problem of learning from confounded expert data in a contextual MDP setup. We analyze the limitations of learning from such data with and without external rewar…

Cited by 19SourcePDFScholar
2022

Reinforcement Learning with a Terminator

NeurIPS 2022accept

We present the problem of reinforcement learning with exogenous termination. We define the Termination Markov Decision Process (TerMDP), an extension of the MDP framework, in which episodes may be interrupted by an external non-Markovian observer. This formulation accounts for numerous real-world si…

2021

Improve Agents without Retraining: Parallel Tree Search with Off-Policy Correction

NeurIPS 2021poster

Tree Search (TS) is crucial to some of the most influential successes in reinforcement learning. Here, we tackle two major challenges with TS that limit its usability: \textit{distribution shift} and \textit{scalability}. We first discover and analyze a counter-intuitive phenomenon: action selection…

2015

Off-policy Model-based Learning under Unknown Factored Dynamics

ICML 2015poster

Off-policy learning in dynamic decision problems is essential for providing strong evidence that a new policy is better than the one in use. But how can we prove superiority without testing the new policy? To answer this question, we introduce the G-SCOPE algorithm that evaluates a new policy based…

Cited by 41SourcePDFScholar