← Search

Jiaxi Gu

8 accepted papers

2026

Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping

ICML 2026poster

Precise camera pose control is critical for video diffusion, yet maintaining geometric consistency remains a challenge. Existing methods that directly inject numerical camera parameters into the diffusion backbone often fail to bridge the gap between abstract coordinates and visual content, leading …

Cited by 1SourceScholar
2025

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

ICASSP 2025accepted

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods frequently integrate semantic information from images or simply concatenate images, which often leads to low fidelity and flickering in the generated videos. T…

Cited by 25SourceScholar
2025

EasyControl: Adding Control to Video Diffusion for Controllable Video Generation and Interpolation

ICASSP 2025accepted

The diffusion model is widely leveraged for either controllable video generation or video interpolation. As each field has its task-specific problems, it is difficult to merely develop a single model for completing both tasks simultaneously. Moreover, most existing works only support image condition…

Cited by 0SourceScholar
2025

MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

ICML 2025poster

In recent years, while generative AI has advanced significantly in image generation, video generation continues to face challenges in controllability, length, and detail quality, which hinder its application. We present MimicMotion, a framework for generating high-quality human videos of arbitrary l…

2024

BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models

CVPR 2024poster

Diffusion models have made tremendous progress in text-driven image and video generation. Now text-to-image foundation models are widely applied to various downstream image synthesis tasks such as controllable image generation and image editing while downstream video synthesis tasks are less explore…

2024

MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing

ECCV 2024poster

"The diffusion model is widely leveraged for either video generation or video editing. As each field has its task-specific problems, it is difficult to merely develop a single diffusion for completing both tasks simultaneously. Video diffusion sorely relying on the text prompt can be adapted to unif…

2023

PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval

ICCV 2023poster

Text-video retrieval is a fundamental task with high practical value in multi-modal research. Inspired by the great success of pre-trained image-text models with large-scale data, such as CLIP, many methods are proposed to transfer the strong representation learning capability of CLIP to text-video…

Cited by 19PDFScholar
2022

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

NeurIPS 2022accept

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks in Chinese hinders the development of Chinese VLP models an…