← Search

Peiran Wu

1 accepted papers

2026

MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding

ICLR 2026poster

The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. Nevertheless, visual language models (VLMs) still face significant computational overhead when scaled from images to the video domain. When video data is too large (due to high frame rates and long dura…

Cited by 0SourceScholar