2025
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
ICLR 2025poster
We present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos that align seamlessly with text prompts, with a frame rate of 16 fps and resolution of 768 x 1360 pixels. Previous video generation models often struggle…