2026
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
CVPR 2026
We introduce TimeViper, a hybrid vision-language model designed to tackle challenges of long video understanding. Processing long videos demands both an efficient model architecture and an effective mechanism for handling extended temporal contexts. To this end, TimeViper adopts a hybrid Mamba-Trans