← Search

Thibaud Rahier

6 accepted papers

2024

Maximizing the Success Probability of Policy Allocations in Online Systems

AAAI 2024technical

The effectiveness of advertising in e-commerce largely depends on the ability of merchants to bid on and win impressions for their targeted users. The bidding procedure is highly complex due to various factors such as market competition, user behavior, and the diverse objectives of advertisers. In t…

2024

Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits

NeurIPS 2024poster

We address the problem of stochastic combinatorial semi-bandits, where a player selects among $P$ actions from the power set of a set containing $d$ base items. Adaptivity to the problem's structure is essential in order to obtain optimal regret upper bounds. As estimating the coefficients of a cova…

Cited by 1SourcePDFScholar
2022

Diverse Weight Averaging for Out-of-Distribution Generalization

NeurIPS 2022accept

Standard neural networks struggle to generalize under distribution shifts in computer vision. Fortunately, combining multiple networks can consistently improve out-of-distribution generalization. In particular, weight averaging (WA) strategies were shown to perform best on the competitive DomainBed…

2021

Zeroth-Order Non-Convex Learning via Hierarchical Dual Averaging

ICML 2021spotlight

We propose a hierarchical version of dual averaging for zeroth-order online non-convex optimization {–} i.e., learning processes where, at each stage, the optimizer is facing an unknown non-convex loss function and only receives the incurred loss as feedback. The proposed class of policies relies on…

Cited by 16SourcePDFScholar
2020

Online Non-Convex Optimization with Imperfect Feedback

NeurIPS 2020poster

We consider the problem of online learning with non-convex losses. In terms of feedback, we assume that the learner observes – or otherwise constructs – an inexact model for the loss function encountered at each stage, and we propose a mixed-strategy learning policy based on dual averaging. In this…

Cited by 26SourcePDFScholar