2024
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
ECCV 2024poster
"Recent vision-language pre-training models have exhibited remarkable generalization ability in zero-shot recognition tasks. Previous open-vocabulary 3D scene understanding methods mostly focus on training 3D models using either image or text supervision while neglecting the collective strength of a…