← Search

Luke Schultz

1 accepted papers

2026

RePack then Refine: Efficient Diffusion Transformers with Vision Foundation Models

ICML 2026poster

Semantic-rich features from Vision Foundation Models (VFMs) have been leveraged to enhance Latent Diffusion Models (LDMs). However, raw VFM features are typically high-dimensional and redundant, increasing the difficulty of learning and reducing training efficiency for Diffusion Transformers (DiTs).…

Cited by 0SourceScholar