2025
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
ICCV 2025poster
We introduce MUSE-VL, a Unified Vision-Language Model through Semantic discrete Encoding for multimodal understanding and generation. Recently, the research community has begun exploring unified models for visual generation and understanding. However, existing vision tokenizers (e.g., VQGAN) only co…