← Search

Mengyi Liu

3 accepted papers

2026

Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding

CVPR 2026

Despite rapid developments and widespread applications of MLLM agents, they still struggle with long-form video understanding (LVU) tasks, which are characterized by high information density and extended temporal spans. Recent research on LVU agents demonstrates that simple task decomposition and co

Cited by 0SourcecodeScholar
2025

Re-Attentional Controllable Video Diffusion Editing

AAAI 2025technical

Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploited large-scale text-to-image diffusion models for text-guided video editing, resu…

2025

STDD: Spatio-Temporal Dual Diffusion for Video Generation

CVPR 2025poster

Diffusion probabilistic model is becoming the cornerstone of data generation, especially generating high-quality images. As an extension, video diffusion generation is in urgent need of a principled temporal-sequence diffusion way, while the spatial-domain diffusion dominates most video diffusion me…

Cited by 0SourcePDFScholar