MMG-VL: A Vision-Language Driven Approach for Multi-Person Motion Generation
Generating realistic and coordinated 3D human motion for multiple individuals within complex environments remains a significant challenge. Existing text-to-motion methods are often ``blind