2026
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
ICLR 2026poster
The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressive paradigm of large language models. A vision tokenizer trained for reconstruction excels at capturing low-level visual appearance, making it well-suited for…