← Search

Gene Li

5 accepted papers

2026

Learning to Answer from Correct Demonstrations

ICLR 2026poster

We study the problem of learning to generate an answer (or completion) to a question (or prompt), where there could be multiple correct answers, any one of which is acceptable at test time. Learning is based on demonstrations of some correct answer to each training question, as in Supervised Fine Tu…

Cited by 0SourceScholar
2023

When is Agnostic Reinforcement Learning Statistically Tractable?

NeurIPS 2023poster

We study the problem of agnostic PAC reinforcement learning (RL): given a policy class $\Pi$, how many rounds of interaction with an unknown MDP (with a potentially large state and action space) are required to learn an $\epsilon$-suboptimal policy with respect to \(\Pi\)? Towards that end, we intro…

Cited by 7SourcePDFScholar
2022

Exponential Family Model-Based Reinforcement Learning via Score Matching

NeurIPS 2022accept

We propose an optimistic model-based algorithm, dubbed SMRL, for finite-horizon episodic reinforcement learning (RL) when the transition model is specified by exponential family distributions with $d$ parameters and the reward is bounded and known. SMRL uses score matching, an unnormalized density e…

2022

Pessimism for Offline Linear Contextual Bandits using $\ell_p$ Confidence Sets

NeurIPS 2022accept

We present a family $\{\widehat{\pi}_p\}_{p\ge 1}$ of pessimistic learning rules for offline learning of linear contextual bandits, relying on confidence sets with respect to different $\ell_p$ norms, where $\widehat{\pi}_2$ corresponds to Bellman-consistent pessimism (BCP), while $\widehat{\pi}_\in…

Cited by 27SourcePDFScholar