2026
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
ICML 2026poster
Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this AR paradigm inevitably faces a dual efficiency bottleneck: strictly unidirectional attention compromises *understanding …