2026
UM3D: Towards a Unified Multimodal 3D Shape Generation Model
ICRA 2026poster
Vision-Language Pre-training models (VLMs) have emerged as a highly promising solution to the generative problem, achieving remarkable success in the field of 2D image generation. However, extending these 2D paradigms to 3D domains is still unexplored due to the scarcity of text-3D pairs and shape a…