2026
Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment
ICML 2026spotlight
We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions. In this setting, aligning style with high task performance is particularly challenging due to distribution shift and inherent conflicts between style and rewar…