← Search

Imad Aouali

8 accepted papers

2025

Bayesian Off-Policy Evaluation and Learning for Large Action Spaces

AISTATS 2025poster

In interactive systems, actions are often correlated, presenting an opportunity for more sample-efficient off-policy evaluation (OPE) and learning (OPL) in large action spaces. We introduce a unified Bayesian framework to capture these correlations through structured and informative priors. In this…

Cited by 0SourceScholar
2025

Diffusion Models Meet Contextual Bandits

NeurIPS 2025poster

Efficient online decision-making in contextual bandits is challenging, as methods without informative priors often suffer from computational or statistical inefficiencies. In this work, we leverage pre-trained diffusion models as expressive priors to capture complex action dependencies and develop a…

Cited by 0SourceScholar
2025

Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits

AISTATS 2025poster

We study the problem of Bayesian fixed-budget best-arm identification (BAI) in structured bandits. We propose an algorithm that uses fixed allocations based on the prior information and the structure of the environment. We provide theoretical bounds on its performance across diverse models, includin…

Cited by 0SourcecodeScholar
2024

Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning

NeurIPS 2024spotlight

This work investigates the offline formulation of the contextual bandit problem, where the goal is to leverage past interactions collected under a behavior policy to evaluate, select, and learn new, potentially better-performing, policies. Motivated by critical applications, we move beyond point est…

2024

Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling

UAI 2024poster

Off-policy learning (OPL) often involves minimizing a risk estimator based on importance weighting to correct bias from the logging policy used to collect data. However, this method can produce an estimator with a high variance. A common solution is to regularize the importance weights and learn the…

Cited by 1SourcePDFScholar