2025
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
ICLR 2025oral
The trustworthiness of Large Language Models (LLMs) refers to the extent to which their outputs are reliable, safe, and ethically aligned, and it has become a crucial consideration alongside their cognitive performance. In practice, Reinforcement Learning From Human Feedback (RLHF) has been widely u…