2026
Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
ICML 2026oral
Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal reasoning tasks; however, their learning remains constrained by the limitations of human-annotated supervision. Recent self-rewarding approaches attempt to overcome this constraint by allowing models to act as their…