2025
MotionBind: Multi-Modal Human Motion Alignment for Retrieval, Recognition, and Generation
NeurIPS 2025poster
Recent advances in multi-modal representation learning have led to unified embedding spaces that align modalities such as images, text, audio, and vision. However, human motion sequences, a modality that is fundamental for understanding dynamic human activities, remains largely unrepresented in thes…