AAAI 2026technical0 citations

Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models

Yucheng Zhou, Jihai Zhang, Guanjie Chen, Jianbing Shen, Yu Cheng

Abstract

Video generation using Large Language Models (LLMs) has shown promising potential, effectively leveraging the extensive LLM infrastructure to provide a unified framework for multimodal understanding and content generation. However, these methods face critical challenges, i.e., token redundancy and inefficiencies arising from long sequences, which constrain their performance and efficiency compared to diffusion-based approaches. In this study, we investigate the impact of token redundancy in LLM-based video generation by information-theoretic analysis and propose Vision Representation Compression (VRC), a novel framework designed to achieve more in both performance and efficiency with less video token representations. VRC introduces learnable representation compressor and decompressor to compress video token representations, enabling autoregressive next-sequence prediction in a compact latent space. Our approach reduces redundancy, shortens token sequences, and improves model

BibTeX
@inproceedings{aaai2026_lessismorevision,
  title = {Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models},
  author = {Yucheng Zhou and Jihai Zhang and Guanjie Chen and Jianbing Shen and Yu Cheng},
  booktitle = {AAAI 2026},
  year = {2026}
}
Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models · AAAI 2026