2026
Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
CVPR 2026
Spatial intelligence in vision-language models (VLMs) attracts research interest with the practical demand to reason in the 3D world. Despite promising results, most existing methods follow the conventional 2D pipeline in VLMs and use pixel-aligned representations for the vision modality. However, c