DynaMem: Consistent Long Video Generation via Hierarchical Memory and Motion Priors
Recent text-to-video diffusion models can synthesize visually compelling clips from natural language prompts. However, practical applications increasingly demand long-form videos with evolving narratives and persistent identity. A common solution is autoregressive generation, where the video is prod…