← Search

Shunsi Zhang

9 accepted papers

2026

Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback

AAAI 2026technical

Recent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, existing approaches still face challenges in lip-sync accuracy, temporal coherence for long video generation, and multi-ch

Cited by 0SourcePDFScholar
2025

FlexGen: Flexible Multi-View Generation from Text and Image Inputs

ICCV 2025poster

In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable multi-view synthesis through additional conditioning on 3D-aware…

Cited by 0SourcePDFScholar
2025

GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs

ICCV 2025poster

Estimating physical properties for visual data is a crucial task in computer vision, graphics, and robotics, underpinning applications such as augmented reality, physical simulation, and robotic grasping. However, this area remains under-explored due to the inherent ambiguities in physical property…

Cited by 0SourcePDFScholar
2025

Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation

CVPR 2025poster

Diffusion models have achieved great success in generating 2D images. However, the quality and generalizability of 3D content generation remain limited. State-of-the-art methods often require large-scale 3D assets for training, which are challenging to collect. In this work, we introduce Kiss3DGen (…

2025

MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

ICLR 2025poster

The recent large-scale text-to-speech (TTS) systems are usually grouped as autoregressive and non-autoregressive systems. The autoregressive systems implicitly model duration but exhibit certain deficiencies in robustness and lack of duration controllability. Non-autoregressive systems require expli…

2025

MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction

CVPR 2025poster

This paper investigates the research task of reconstructing the 3D clothed human body from a monocular image. Due to the inherent ambiguity of single-view input, existing approaches leverage pre-trained SMPL(-X) estimation models or generative models to provide auxiliary information for human recons…

Cited by 3SourcePDFScholar
2025

Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion

ICML 2025poster

Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter significant challenges due to uncontrollable factors, such as in…

2025

Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion

CVPR 2025poster

Rendering and inverse rendering are pivotal tasks in both computer vision and graphics. The rendering equation is the core of the two tasks, as an ideal conditional distribution transfer function from intrinsic properties to RGB images. Despite achieving promising results of existing rendering metho…

Cited by 0SourcePDFScholar
2024

Adaptive-Avg-Pooling Based Attention Vision Transformer for Face Anti-Spoofing

ICASSP 2024accepted

Traditional vision transformer consists of two parts: transformer encoder and multi-layer perception (MLP). The former plays the role of feature learning to obtain better representation, while the latter plays the role of classification. Here, the MLP is constituted of two fully connected (FC) layer…

Cited by 0SourceScholar