2025
LLM Safety Alignment is Divergence Estimation in Disguise
NeurIPS 2025poster
We present a theoretical framework showing that popular LLM alignment methods—including RLHF and its variants—can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less-preferred) distributions. This perspective explains the emergence of separation…