2026
PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
CVPR 2026
The development of 3D Vision-Language Models (VLMs), crucial for applications in robotics, autonomous driving, and augmented reality, is severely constrained by the scarcity of paired 3D-text data. Existing methods rely solely on next-token prediction loss, using only language tokens for supervision