2025
Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
ICCV 2025poster
We present a self-supervised method to improve an agent's abilities in describing arbitrary objects while actively exploring a generic environment. This is a challenging problem, as current models struggle to obtain coherent image captions due to different camera viewpoints and clutter. We propose a…