Reasoning Diffusion for Unpaired Test Time Out-of-distribution Text-Image to Video Generation
Text-image to video generation aims to synthesize a video conditioned on the given text-image inputs. Nevertheless, existing methods generally assume that the semantic information carried in the input text and image tends to be perfectly paired and temporally aligned, occurring simultaneously in the