2025
Mamba Fusion: Learning Actions Through Questioning
ICASSP 2025accepted
Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning. While transformer-based architectures have been the de facto in vision-language training, they face challenges like quadratic computational complexity, high GPU memory usage, a…