← Search

Wanqiao Xu

3 accepted papers

2026

Formalizing Learning from Language Feedback with Provable Guarantees

ICML 2026poster

Interactively learning from observation and language feedback is an increasingly studied area driven by the emergence of large language model (LLM) agents. While impressive empirical demonstrations have been shown, so far a principled framing of these decision problems remains lacking. In this paper…

Cited by 0SourceScholar
2023

Uniformly Conservative Exploration in Reinforcement Learning

AISTATS 2023poster

A key challenge to deploying reinforcement learning in practice is avoiding excessive (harmful) exploration in individual episodes. We propose a natural constraint on exploration—uniformly outperforming a conservative policy (adaptively estimated from all data observed thus far), up to a per-episode…

Cited by 6SourcePDFScholar
2022

Regret Bounds for Risk-Sensitive Reinforcement Learning

NeurIPS 2022accept

In safety-critical applications of reinforcement learning such as healthcare and robotics, it is often desirable to optimize risk-sensitive objectives that account for tail outcomes rather than expected reward. We prove the first regret bounds for reinforcement learning under a general class of risk…

Cited by 23SourcePDFScholar