← Search

Claire Vernade

16 accepted papers

2025

Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits

AISTATS 2025poster

We study the problem of Bayesian fixed-budget best-arm identification (BAI) in structured bandits. We propose an algorithm that uses fixed allocations based on the prior information and the structure of the environment. We provide theoretical bounds on its performance across diverse models, includin…

Cited by 0SourcecodeScholar
2025

Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning

NeurIPS 2025poster

The Combined Algorithm Selection and Hyperparameter optimization (CASH) is a challenging resource allocation problem in the field of AutoML. We propose MaxUCB, a max $k$-armed bandit method to trade off exploring different model classes and conducting hyperparameter optimization. MaxUCB is specific…

Cited by 0SourcecodeScholar
2025

Quantization-Free Autoregressive Action Transformer

NeurIPS 2025spotlight

Current transformer-based imitation learning approaches introduce discrete action representations and train an autoregressive transformer decoder on the resulting latent code. However, the initial quantization breaks the continuous structure of the action space thereby limiting the capabilities of t…

Cited by 0SourcecodeScholar
2022

EigenGame Unloaded: When playing games is better than optimizing

ICLR 2022poster

We build on the recently proposed EigenGame that views eigendecomposition as a competitive game. EigenGame's updates are biased if computed using minibatches of data, which hinders convergence and more sophisticated parallelism in the stochastic setting. In this work, we propose an unbiased stochast…

Cited by 12SourcePDFScholar
2021

Confident Off-Policy Evaluation and Selection through Self-Normalized Importance Weighting

AISTATS 2021poster

We consider off-policy evaluation in the contextual bandit setting for the purpose of obtaining a robust off-policy selection strategy, where the selection strategy is evaluated based on the value of the chosen policy in a set of proposal (target) policies. We propose a new method to compute a lower…

2020

Linear bandits with Stochastic Delayed Feedback

ICML 2020poster

Stochastic linear bandits are a natural and well-studied model for structured exploration/exploitation problems and are widely used in applications such as on-line marketing and recommendation. One of the main challenges faced by practitioners hoping to apply existing algorithms is that usually the…

Cited by 88SourcePDFScholar
2020

Stochastic bandits with arm-dependent delays

ICML 2020poster

Significant work has been recently dedicated to the stochastic delayed bandits because of its relevance in applications. The applicability of existing algorithms is however restricted by the fact that strong assumptions are often made on the delay distributions, such as full observability, restricti…

Cited by 66SourcePDFScholar