LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
Despite achieving significant progress in 2D image understanding, large multimodal models (LMMs) struggle in the physical world due to the lack of spatial representation. Typically, existing 3D LMMs mainly embed 3D positions as fixed spatial prompts within visual features to represent the scene. How…