2025
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
EMNLP 2025
Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain fundamentally limited in their understanding of 3D spatial structures.We propose Geometric Distillation, a lightweight, annotation-free fine-tuning framework that injects human-ins