2026
TV2TV: A Unified Framework for Interleaved Language and Video Generation
CVPR 2026
Video generation models are rapidly advancing, but can still struggle with complex video outputs that require significant semantic branching or repeated high-level reasoning about what should happen next. In this paper, we introduce a new class of omni video-text models that integrate ideas from rec