2025
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
ICLR 2025poster
As LLMs become more widely deployed, there is increasing interest in directly optimizing for feedback from end users (e.g. thumbs up) in addition to feedback from paid annotators. However, training to maximize human feedback creates a perverse incentive structure for the AI to resort to manipulative…