← Search

Otmane Sakhi

4 accepted papers

2024

Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning

NeurIPS 2024spotlight

This work investigates the offline formulation of the contextual bandit problem, where the goal is to leverage past interactions collected under a behavior policy to evaluate, select, and learn new, potentially better-performing, policies. Motivated by critical applications, we move beyond point est…

2023

Fast Offline Policy Optimization for Large Scale Recommendation

AAAI 2023technical

Personalised interactive systems such as recommender systems require selecting relevant items from massive catalogs dependent on context. Reward-driven offline optimisation of these systems can be achieved by a relaxation of the discrete problem resulting in policy learning or REINFORCE style learni…