← Search

Paras Chopra

2 accepted papers

2025

IPO: Your Language Model is Secretly a Preference Classifier

ACL 2025long

Reinforcement learning from human feedback (RLHF) has emerged as the primary method for aligning large language models (LLMs) with human preferences. While it enables LLMs to achieve human-level alignment, it often incurs significant computational and financial costs due to its reliance on training…