2025
Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment
AAAI 2025technical
The alignment of large language models (LLMs) is crucial for generating helpful and harmless content. Existing approaches leverage preference-based human feedback data to learn the reward function and align the LLM with the feedback data. However, these approaches focus on modeling the reward differ…