DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
Reducing the token count is crucial for efficient training and inference of latent diffusion models, especially at high resolution. A common approach is to build high-compression image tokenizers that store more information by allocating more channels per token. However, when trained solely with rec