2026
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision–Language Models
ICLR 2026poster
Vision–Language Models (VLMs) show promise as zero-shot goal-conditioned value functions, but their frozen pre-trained representations limit generalization and temporal reasoning. We introduce VITA, a zero-shot value function learning method that enhances both capabilities via test-time adaptation.…