← Search

Mingzhen Sun

8 accepted papers

2026

VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis

ICLR 2026poster

The notable gap between user-provided and model-preferred prompts poses a significant challenge for generating high-quality images with text-to-image models, compelling the need for prompt engineering. Current studies on prompt engineering can effectively enhance the style and aesthetics of generate…

Cited by 0SourcecodeScholar
2026

W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image Editing

ICLR 2026poster

While recent advances in Diffusion Transformers (DiTs) have significantly advanced text-to-image generation, text-driven image editing remains challenging. Existing approaches either struggle to balance structural preservation with flexible modifications or require costly fine-tuning of large models…

Cited by 0SourceScholar
2025

AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion

CVPR 2025poster

The task of video generation requires synthesizing visually realistic and temporally coherent video frames. Existing methods primarily use asynchronous auto-regressive models or synchronous diffusion models to address this challenge. However, asynchronous auto-regressive models often suffer from inc…

2025

I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength

ICLR 2025poster

Video generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However, existing camera control methods still suffer from several limita…

2025

I2VControl: Disentangled and Unified Video Motion Synthesis Control

ICCV 2025poster

Motion controllability is crucial in video synthesis. However, most previous methods are limited to single control types, and combining them often results in logical conflicts. In this paper, we propose a disentangled and unified framework, namely I2VControl, to overcome the logical conflicts. We re…

2023

GLOBER: Coherent Non-autoregressive Video Generation via GLOBal Guided Video DecodER

NeurIPS 2023poster

Video generation necessitates both global coherence and local realism. This work presents a novel non-autoregressive method GLOBER, which first generates global features to obtain comprehensive global guidance and then synthesizes video frames based on the global features to generate coherent videos…

2023

MOSO: Decomposing MOtion, Scene and Object for Video Prediction

CVPR 2023poster

Motion, scene and object are three primary visual components of a video. In particular, objects represent the foreground, scenes represent the background, and motion traces their dynamics. Based on this insight, we propose a two-stage MOtion, Scene and Object decomposition framework (MOSO) for video…

2023

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

NeurIPS 2023poster

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections between multi-modality video tracks, including Vision, Audi…