Zero-Shot Text-to-Motion Evaluation using Video Language Models
Text-to-motion (T2M) generation has emerged as a fundamental task. However, existing evaluation metrics often fail to accurately capture the semantic alignment between textual descriptions and generated 3D motions. In this work, we propose VeMo, a novel evaluation framework that leverages the zero-s…