← Search

Pierre Clavier

5 accepted papers

2025

ShiQ: Bringing back Bellman to LLMs

NeurIPS 2025poster

The fine-tuning of pre-trained large language models (LLMs) using reinforcement learning (RL) is generally formulated as direct policy optimization. This approach was naturally favored as it efficiently improves a pretrained LLM with simple gradient updates. Another RL paradigm, Q-learning methods,…

Cited by 3SourceScholar
2024

$\mathtt{VITS}$ : Variational Inference Thompson Sampling for contextual bandits

ICML 2024poster

In this paper, we introduce and analyze a variant of the Thompson sampling (TS) algorithm for contextual bandits. At each round, traditional TS requires samples from the current posterior distribution, which is usually intractable. To circumvent this issue, approximate inference techniques can be us…

Cited by 3SourcePDFScholar
2024

Near-Optimal Distributionally Robust Reinforcement Learning with General $L_p$ Norms

NeurIPS 2024poster

To address the challenges of sim-to-real gap and sample efficiency in reinforcement learning (RL), this work studies distributionally robust Markov decision processes (RMDPs) --- optimize the worst-case performance when the deployed environment is within an uncertainty set around some nominal MDP. D…

Cited by 0SourcePDFScholar