← Search

Rajdeep Haldar

2 accepted papers

2025

LLM Safety Alignment is Divergence Estimation in Disguise

NeurIPS 2025poster

We present a theoretical framework showing that popular LLM alignment methods—including RLHF and its variants—can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less-preferred) distributions. This perspective explains the emergence of separation…

Cited by 0SourcecodeScholar