2026
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
ICLR 2026poster
Diffusion models have shown impressive performance in many visual generation and manipulation tasks. Many existing methods focus on training a model for a specific task, especially, text-to-video (T2V) generation, while many other works focus on finetuning the pretrained T2V model for image-to-video…