← Search

Zhan Xu

11 accepted papers

2026

AnchorSplat: Feed-Forward 3D Gaussian Splatting With 3D Geometric Priors

CVPR 2026

Recent feed-forward Gaussian reconstruction models adopt a pixel-aligned formulation that maps each 2D pixel to a 3D Gaussian, entangling Gaussian representations tightly with the input images. In this paper, we propose AnchorSplat, a novel feed-forward 3DGS framework for scene-level reconstruction

Cited by 0SourceScholar
2026

SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenes

CVPR 2026

We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling, trained solely on differentiable renderings without any flow supervision. Besi

Cited by 0SourcecodeScholar
2025

Free-viewpoint Human Animation with Pose-correlated Reference Selection

CVPR 2025highlight

Diffusion-based human animation aims to animate a human character based on a source human image as well as driving signals such as a sequence of poses. Leveraging the generative capacity of diffusion model, existing approaches are able to generate high-fidelity poses, but struggle with significant v…

Cited by 1SourcePDFScholar
2025

Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface Networks

ICCV 2025poster

Diffusion Transformers (DiTs) can generate short photorealistic videos, yet directly training and sampling longer videos with full attention across the video remains computationally challenging. Alternative methods break long videos down into sequential generation of short video segments, requiring…

2025

Move-in-2D: 2D-Conditioned Human Motion Generation

CVPR 2025poster

Generating realistic human videos remains a challenging task, with the most effective methods currently relying on a human motion sequence as a control signal. Existing approaches often use existing motion extracted from other videos, which restricts applications to specific motion types and global…

2025

Visual Persona: Foundation Model for Full-Body Human Customization

CVPR 2025poster

We introduce Visual Persona, a foundation model for text-to-image full-body human customization that, given a single in-the-wild human image, generates diverse images of the individual guided by text descriptions. Unlike prior methods that focus solely on preserving facial identity, our approach cap…

Cited by 0SourcePDFScholar
2024

ActAnywhere: Subject-Aware Video Background Generation

NeurIPS 2024poster

We study a novel problem to automatically generate video background that tailors to foreground subject motion. It is an important problem for the movie industry and visual effects community, which traditionally requires tedious manual efforts to solve. To this end, we propose ActAnywhere, a video di…

2022

APES: Articulated Part Extraction From Sprite Sheets

CVPR 2022poster

Rigged puppets are one of the most prevalent representations to create 2D character animations. Creating these puppets requires partitioning characters into independently moving parts. In this work, we present a method to automatically identify such articulated parts from a small set of character po…

Cited by 5PDFcodeScholar
2020

Contextual Residual Aggregation for Ultra High-Resolution Image Inpainting

CVPR 2020oral

Recently data-driven image inpainting methods have made inspiring progress, impacting fundamental image editing tasks such as object removal and damaged image repairing. These methods are more effective than classic approaches, however, due to memory limitations they can only handle low-resolution i…

Cited by 443PDFcodeScholar