2025
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
ICCV 2025poster
We introduce DIP, a novel unsupervised post-training method designed to enhance dense representations in large-scale pretrained vision encoders for in-context scene understanding. Unlike prior approaches using complex self-distillation architectures, our method trains the vision encoder using pseudo…