← Search

Victoriano Montesinos

2 accepted papers

2026

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

RSS 2026poster

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy must implicitly infer complex physical dynamics and tempora…

Cited by 0SourceScholar
2024

Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

ICLR 2024poster

Reinforcement learning (RL) requires either manually specifying a reward function, which is often infeasible, or learning a reward model from a large amount of human feedback, which is often very expensive. We study a more sample-efficient alternative: using pretrained vision-language models (VLMs)…