2026
MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
ICASSP 2026poster
Multi-image reasoning and grounding require understanding complex cross-image relationships at both object levels and image levels. Current Large Visual Language Models (LVLMs) face two critical challenges: the lack of cross-image reasoning capabilities and insufficient cross-image reference reward…