2026
ROSETTA: Constructing Code-Based Reward from Unconstrained Language Preference
ICLR 2026poster
Intelligent embodied agents not only need to accomplish preset tasks, but also learn to align with individual human needs and preferences. Extracting reward signals from human language preferences allows an embodied agent to adapt through reinforcement learning. However, human language preferences a…