2026
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
CVPR 2026
Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. We introduce NUMINA, a training-free identify-then-guide framework for improved numerical alignment. NUMINA identifies prompt-layout inconsi