ICML 2026poster0 citations

Vision in One Vector: Implicit Visual Compression with Diffusion Foundation Models

Jiajun He, Zongyu Guo, Zhaoyang Jia, Xiaoyi Zhang, Jiahao Li, Xiao Li, Bin Li, Jose Miguel Hernandez-Lobato

Abstract

Modern visual generative models acquire rich visual knowledge through large-scale training, yet existing visual representations (such as pixels, latents, or tokens) remain external to the model and cannot directly exploit this knowledge for compact storage or reuse. In this work, we introduce a new visual representation framework that encodes a signal as a function, which is parametrized by low-rank adaptations attached to a frozen visual generative model. Such implicit representations are learned by compressing the visual signal, e.g., an 81-frame video, into a single compact vector, achieving strong perceptual video compression at extremely low bitrates. Beyond basic compression, the functional nature of this representation enables inference-time scaling and control, allowing additional refinement on the compression performance. More broadly, as the implicit representations directly act as a function of the generation process, this suggests a unified framework bridging visual compression and generation.

DiffusionVisionRetrieval
BibTeX
@inproceedings{
guo2026compression,
title={Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation Models},
author={Zongyu Guo and Jiajun He and Zhaoyang Jia and Xiaoyi Zhang and Jiahao Li and Xiao Li and Bin Li and Jos{\'e} Miguel Hern{\'a}ndez-Lobato and Yan Lu},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=eYICUJ62hg}
}