2025
Differentiable Hierarchical Visual Tokenization
NeurIPS 2025spotlight
Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images. In this work, we introduce an end-to-end differentiable tokenizer that adapts to image content with pixel-level granularity while remaining backward-compatible with existing architectures for ret…