← Search

Maziar Raissi

1 accepted papers

2025

Aligning to What? Limits to RLHF Based Alignment

NAACL 2025findings

Reinforcement Learning from Human Feedback (RLHF) is increasingly used to align large language models (LLMs) with human preferences. However, the effectiveness of RLHF in addressing underlying biases remains unclear. This study investigates the relationship between RLHF and both covert and overt bia…