2026
Reducing Oracle Feedback with Vision Language Embeddings for Preference Based RL
ICRA 2026poster
Preference-based reinforcement learning (RL) offers a promising approach for aligning policies with human intent but is often constrained by the high cost of human feedback. In this work, we introduce ROVED, a framework that integrates Vision-Language Models (VLMs) with selective human feedback to s…