Position: Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences
This position paper argues that \textbf{large language models should transition from learning aggregated human preferences to learning personalized, individual preferences}. Current approaches to training language models with reinforcement learning from human feedback (RLHF) aggregate diverse human …