← Search

Zhensong Zhang

20 accepted papers

2026

CHROMA: Consistent Harmonization of Multi-View Appearance via Bilateral Grid Prediction

ICLR 2026poster

Modern camera pipelines apply extensive on-device processing, such as exposure adjustment, white balance, and color correction, which, while beneficial individually, often introduce photometric inconsistencies across views. These appearance variations violate multi-view consistency and degrade novel…

Cited by 0SourceScholar
2026

Charge: A Comprehensive Novel View Synthesis Benchmark and Dataset to Bind Them All

CVPR 2026

This paper presents a new dataset for Novel View Synthesis, generated from a high-quality, animated film with stunning realism and intricate detail. Our dataset captures a variety of dynamic scenes, complete with detailed textures, lighting, and motion, making it ideal for training and evaluating cu

Cited by 0SourceScholar
2026

Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing

CVPR 2026

Always-on sensing is essential for next-generation edge/wearable AI systems, yet continuous high-fidelity RGB video capture remains prohibitively expensive for resource-constrained mobile and edge platforms. We present a new paradigm for efficient streaming video understanding: grayscale-always, col

Cited by 0SourcecodeScholar
2026

Diffusion-Based Makeup Transfer with Facial Region-Aware Makeup Features

CVPR 2026

Current diffusion-based makeup transfer methods commonly use the makeup information encoded by off-the-shelf foundation models (e.g., CLIP) as condition to preserve the makeup style of reference image in the generation. Although effective, these works mainly have two limitations: (1) foundation mode

Cited by 0SourcecodeScholar
2026

LiteVSR: Enabling Cross-Domain Fine-Grained Detail Generation in Light-Weight Transformers for Video Super-Resolution

ICML 2026poster

Large-scale pre-trained video generators offer powerful priors for Video Super-Resolution (VSR), yet adapting them remains computationally prohibitive. Full fine-tuning demands extensive resources, and ControlNet-style adapters lose their efficiency advantage under modern Diffusion Transformers (DiT…

Cited by 0SourceScholar
2026

Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting

CVPR 2026

Feed-forward 3D Gaussian Splatting (3DGS) models enable real-time scene generation but are hindered by suboptimal pixel-aligned primitive placement, which relies on a dense, rigid grid that limits both quality and efficiency. We introduce a new feed-forward architecture that detects 3D Gaussian prim

Cited by 0SourceScholar
2026

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

AAAI 2026technical

The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs)

Cited by 0SourcePDFScholar
2025

CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth

CVPR 2025poster

We present CaricatureBooth, a system that transforms caricature creation into a simple interactive experience -- as easy as using a photo booth! A key challenge in caricature generation is two-fold: the scarcity of high-quality caricature data and the difficulty in enabling precise creative control…

2025

Frequency-Guided Diffusion for Training-Free Text-Driven Image Translation

ICCV 2025poster

Current training-free text-driven image translation primarily uses diffusion features (convolution and attention) of pre-trained model as guidance to preserve the style/structure of source image in translated image. However, the coarse guidance at feature level struggles with style (e.g., visual pat…

Cited by 0SourcePDFScholar
2025

Identity-Preserving Audio-Driven Holistic Human Motion Video Generation

ICASSP 2025accepted

Generating realistic human motion videos is a pivotal challenge in advancing human-computer interaction. While existing approaches often focus on generating either head or gesture movements from audio, they lack unified control over full-body motion, frequently producing low-resolution and blurred o…

Cited by 0SourceScholar
2025

ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs

NeurIPS 2025poster

Dynamic Novel View Synthesis aims to generate photorealistic views of moving subjects from arbitrary viewpoints. This task is particularly challenging when relying on monocular video, where disentangling structure from motion is ill-posed and supervision is scarce. We introduce Video Diffusion-Aware…

Cited by 0SourceScholar
2025

VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweening

IJCAI 2025

We propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar

Cited by 0SourcePDFScholar
2024

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

CVPR 2024poster

Co-speech gestures if presented in the lively form of videos can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons resulting in the omission of appearance information we focus on the direct generation of audio-driven co-spee…

2024

Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models

ICASSP 2024accepted

Audio-driven co-speech human gesture generation has made remarkable advancements recently. However, most previous works only focus on single person audio-driven gesture generation. We aim at solving the problem of conversational co-speech gesture generation that considers multiple participants in a…

Cited by 0SourceScholar
2024

Low-Res Leads the Way: Improving Generalization for Super-Resolution by Self-Supervised Learning

CVPR 2024poster

For image super-resolution (SR) bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel "Low-Res Leads the Way" (LWay) training framework merging Supervised Pre-training with Self-supervised Learning to enh…

Cited by 15SourcePDFScholar
2024

SCRREAM : SCan, Register, REnder And Map: A Framework for Annotating Accurate and Dense 3D Indoor Scenes with a Benchmark

NeurIPS 2024poster

Traditionally, 3d indoor datasets have generally prioritized scale over ground-truth accuracy in order to obtain improved generalization. However, using these datasets to evaluate dense geometry tasks, such as depth rendering, can be problematic as the meshes of the dataset are often incomplete and…

2024

Semantics-aware Motion Retargeting with Vision-Language Models

CVPR 2024poster

Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However most of the previous works neglect the semantic information or rely on human-designed joint-level representations. Here we present a novel Semantics-aware Motion reTargeting (SMT) metho…

Cited by 5SourcePDFScholar
2023

DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models

IJCAI 2023poster

The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the rhythm and semantics of the gesture to the corresponding spee…

2023

QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation

CVPR 2023highlight

Speech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a novel quantization-based and phase-guided motion matching framew…

2022

CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation

ECCV 2022poster

"Top-down methods dominate the field of 3D human pose and shape estimation, because they are decoupled from human detection and allow researchers to focus on the core problem. However, cropping, their first step, discards the location information from the very beginning, which makes themselves unabl…