2025
Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
ICLR 2025poster
Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D…