2026
VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
CVPR 2026
While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information. Unlike humans who naturally bridge details and high-level concepts, models tend to treat these elements in isolation. Prevailing evaluation pr