Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization
Solving complex, long-horizon robotic manipulation tasks requires a deep understanding of physical interactions, reasoning about their long-term consequences, and precise high-level planning. Vision-Language Models (VLMs) offer a general perceive-reason-act framework for this goal. However, previous…