2026
Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming Video
ICLR 2026poster
Multimodal large language models have demonstrated impressive capabilities in visual-language understanding, particularly in offline video tasks. More recently, the emergence of online video modeling has introduced early forms of active interaction. However, existing models, typically limited to ten…