← Search

Lee Spector

2 accepted papers

2024

Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization

ICML 2024poster

Reinforcement Learning from Human Feedback (RLHF) has shown potential in qualitative tasks where easily defined performance measures are lacking. However, there are drawbacks when RLHF is commonly used to optimize for average human preferences, especially in generative tasks that demand diverse mode…