2026
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
CVPR 2026
Multimodal reasoning over long-horizon video is challenging due to the need for precise spatiotemporal fusion and alignment across modalities. While recent methods such as Group Relative Policy Optimization (GRPO) have shown promise in this domain, they suffer from three key limitations: (1) data in