2025
OURO: A Self-Bootstrapped Framework for Enhancing Multimodal Scene Understanding
ICCV 2025poster
Multimodal large models have made significant progress, yet fine-grained understanding of complex scenes remains a challenge. High-quality, large-scale vision-language datasets are essential for addressing this issue. However, existing methods often rely on labor-intensive manual annotations or clos…