HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models
3D understanding has drawn significant attention recently, leveraging Vision-Language Models (VLMs) to enable multi-modal reasoning between point cloud and text data. Current 3D-VLMs directly embed the 3D point clouds into 3D tokens, following large 2D-VLMs with powerful reasoning capabilities. Howe