← Search

Shayda Moezzi

2 accepted papers

2026

MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis

CVPR 2026

While text-to-video (T2V) generation has achieved remarkable progress in photorealism, generating intent-aligned videos that faithfully obey physics principles remains a core challenge. In this work, we systematically study Newtonian motion-controlled text-to-video generation and evaluation, emphasi

Cited by 0SourcecodeScholar
2026

Position: Video LLMs Must Not Ignore the Pixel Dynamics in Plain Sight

ICML 2026poster

The essence of video lies in pixel dynamics: motion, state transitions, and the flow of visual information across frames. Video Large Language Models (LLMs) have rapidly become the dominant paradigm for video understanding in computer vision, sophisticated multimodal reasoning over complex, long-for…

Cited by 0SourceScholar