MOSAIC: Multi-Objective Optimization from Zero-Shot Language Reasoning in Preference-Based RL
Preference-based Reinforcement Learning (RL) enables humans to shape complex goals via preference comparisons between sequences of state-action pairs. Most of the existing approaches focus on a singular objective, overlooking the complex causal reasoning that underpins preferences. However, many rea…