Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning