2025
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
NeurIPS 2025poster
Reinforcement learning (RL) has become a predominant technique to align language models (LMs) with human preferences or promote outputs which are deemed to be desirable by a given reward function. Standard RL approaches optimize average reward, while methods explicitly focused on reducing the probab…