← Search

Stephen Zhao

4 accepted papers

2025

Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference

NeurIPS 2025poster

Reinforcement learning (RL) has become a predominant technique to align language models (LMs) with human preferences or promote outputs which are deemed to be desirable by a given reward function. Standard RL approaches optimize average reward, while methods explicitly focused on reducing the probab…

Cited by 0SourceScholar
2024

Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo

ICML 2024oral

Numerous capability and safety techniques of Large Language Models (LLMs), including RLHF, automated red-teaming, prompt engineering, and infilling, can be cast as sampling from an unnormalized target distribution defined by a given reward or potential function over the full sequence. In this work,…

2022

Proximal Learning With Opponent-Learning Awareness

NeurIPS 2022accept

Learning With Opponent-Learning Awareness (LOLA) (Foerster et al. [2018a]) is a multi-agent reinforcement learning algorithm that typically learns reciprocity-based cooperation in partially competitive environments. However, LOLA often fails to learn such behaviour on more complex policy spaces para…

2020

Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning

ICML 2020poster

What goals should a multi-goal reinforcement learning agent pursue during training in long-horizon tasks? When the desired (test time) goal distribution is too distant to offer a useful learning signal, we argue that the agent should not pursue unobtainable goals. Instead, it should set its own intr…