RSS 2025poster1 citations

From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment

Yilin Wu, Thomas Tian, Gokul Swamy, Andrea Bajcsy

Abstract

While generative robot policies have demonstrated significant potential in learning complex, multimodal behaviors from demonstrations, they still exhibit diverse failures at deployment-time. Policy steering offers an elegant solution to reducing the chance of failure by using an external verifier to select from low-level actions proposed by an imperfect generative policy. Here, one might hope to use a Vision Language Model (VLM) as a verifier, leveraging their open-world reasoning capabilities. However, off-the-shelf VLMs struggle to understand the consequences of low-level robot actions as they are represented fundamentally differently than the text and images the VLM was trained on. In response, we propose FOREWARN, a novel framework to unlock the potential of VLMs as open-vocabulary verifiers for runtime policy steering. Our key idea is to decouple the VLM’s burden of predicting action outcomes (

BibTeX
@inproceedings{rss2025_fromforesighttof,
  title = {From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment},
  author = {Yilin Wu and Thomas Tian and Gokul Swamy and Andrea Bajcsy},
  booktitle = {RSS 2025},
  year = {2025}
}
From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment · RSS 2025