2026
MOVi: Training-free Text-conditioned Multi-Object Video Generation
ICASSP 2026oral
Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objects. Most models struggle with accurately capturing complex object interactions, often treating some objects as static ba…