← Search

Shengping Zhang

31 accepted papers

2026

DialoGen: Towards Dialog Gesture Generation via Identity-Decoupled Style Guidance in Interactive Diffusion Model

AAAI 2026technical

We propose DialoGen, a novel framework for generating realistic gestures for both interlocutors in dialog scenarios, conditioned on conversational audios. Unlike most existing methods that focus solely on a single speaker, DialoGen simultaneously generates synchronized gestures for both participants

Cited by 0SourcePDFScholar
2026

MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On

CVPR 2026

Virtual try-on aims to fit an in-shop clothing image onto a specific human body. An optimal virtual try-on method should provide diverse and flexible dressing options, accurately reflecting the varied wearing styles encountered in real-life scenarios, tailored to individual preferences and fashion a

Cited by 0SourceScholar
2026

QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal Evidence

CVPR 2026

Open-vocabulary 3D object affordance grounding aims to identify functional regions of objects given arbitrary semantic descriptions. However, existing methods often rely on fixed training categories and geometric priors, lacking geometric invariance and analogical reasoning capabilities. Since there

Cited by 0SourceScholar
2026

TapSampling: Inference-Time Sampling with a Task-Progress-Understanding Verifier for Robotic Manipulation

ICML 2026poster

Existing embodied control research demonstrates remarkable performance improvements by scaling training data and model size. We instead explore inference-time strategy as an alternative axis. Non-deterministic generative models, such as diffusion and autoregressive models, have been widely adopted i…

Cited by 0SourceScholar
2025

Multi-view Consistent 3D Panoptic Scene Understanding

AAAI 2025technical

3D panoptic scene understanding seeks to create novel view images with 3D-consistent panoptic segmentation, which is crucial for many vision and robotics applications. Mainstream methods (e.g., Panoptic Lifting) directly use machine-generated 2D panoptic segmentation masks as training labels. Howeve…

Cited by 0SourcePDFScholar
2025

Path-Adaptive Matting for Efficient Inference Under Various Computational Cost Constraints

AAAI 2025technical

In this paper, we explore a novel image matting task aimed at achieving efficient inference under various computational cost constraints, specifically FLOP limitations, using a single matting network. Existing matting methods which have not explored scalable architectures or path-learning strategies…

Cited by 0SourcePDFScholar
2025

ProsodyTalker: 3D Visual Speech Animation via Prosody Decomposition

AAAI 2025technical

Most existing 3D visual speech animation methods synthesize lip movements synchronized with speech, which however neglect head poses and therefore degrade the animation realism. The animation of head poses presents two primary challenges: (1) the intricate mapping between speech and head poses remai…

Cited by 0SourcePDFScholar
2024

Autoregressive Queries for Adaptive Tracking with Spatio-Temporal Transformers

CVPR 2024poster

The rich spatio-temporal information is crucial to capture the complicated target appearance variations in visual tracking. However most top-performing tracking algorithms rely on many hand-crafted components for spatio-temporal information aggregation. Consequently the spatio-temporal information i…

2024

DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-based Human Video Generation

CVPR 2024poster

Existing diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In this paper we propose a novel framework DiffPerformer to synthesize high-fidelity an…

Cited by 1SourcePDFScholar
2024

Explicit Visual Prompts for Visual Object Tracking

AAAI 2024technical

How to effectively exploit spatio-temporal information is crucial to capture target appearance changes in visual tracking. However, most deep learning-based trackers mainly focus on designing a complicated appearance model or template updating strategy, while lacking the exploitation of context betw…

2024

Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring

CVPR 2024poster

Existing joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models they typically rely on the assumed degra…

2024

GPS-Gaussian: Generalizable Pixel-wise 3D Gaussian Splatting for Real-time Human Novel View Synthesis

CVPR 2024highlight

We present a new approach termed GPS-Gaussian for synthesizing novel views of a character in a real-time manner. The proposed method enables 2K-resolution rendering under a sparse-view camera setting. Unlike the original Gaussian Splatting or neural implicit rendering methods that necessitate per-su…

2024

GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D Gaussians

CVPR 2024poster

We present GaussianAvatar an efficient approach to creating realistic human avatars with dynamic 3D appearances from a single video. We start by introducing animatable 3D Gaussians to explicitly represent humans in various poses and clothing styles. Such an explicit and animatable representation can…

2024

High-Resolution Image Harmonization with Adaptive-Interval Color Transformation

NeurIPS 2024poster

Existing high-resolution image harmonization methods typically rely on global color adjustments or the upsampling of parameter maps. However, these methods ignore local variations, leading to inharmonious appearances. To address this problem, we propose an Adaptive-Interval Color Transformation meth…

2024

Learning Scale-Aware Spatio-temporal Implicit Representation for Event-based Motion Deblurring

ICML 2024poster

Existing event-based motion deblurring methods mostly focus on restoring images with the same spatial and temporal scales as events. However, the unknown scales of images and events in the real world pose great challenges and have rarely been explored. To address this gap, we propose a novel Scale-A…

2024

Mind the Boundary: Coreset Selection via Reconstructing the Decision Boundary

ICML 2024poster

Existing paradigms of pushing the state of the art require exponentially more training data in many fields. Coreset selection seeks to mitigate this growing demand by identifying the most efficient subset of training data. In this paper, we delve into geometry-based coreset methods and preliminarily…

Cited by 11SourcePDFScholar
2024

ODTrack: Online Dense Temporal Token Learning for Visual Tracking

AAAI 2024technical

Online contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, t…

2024

ProxyCap: Real-time Monocular Full-body Capture in World Space via Human-Centric Proxy-to-Motion Learning

CVPR 2024poster

Learning-based approaches to monocular motion capture have recently shown promising results by learning to regress in a data-driven manner. However due to the challenges in data collection and network designs it remains challenging to achieve real-time full-body capture while being accurate in world…

Cited by 13SourcePDFScholar
2024

Rethinking Imbalance in Image Super-Resolution for Efficient Inference

NeurIPS 2024poster

Existing super-resolution (SR) methods optimize all model weights equally using $\mathcal{L}_1$ or $\mathcal{L}_2$ losses by uniformly sampling image patches without considering dataset imbalances or parameter redundancy, which limits their performance. To address this, we formulate the image SR tas…

Cited by 0SourcePDFScholar
2024

Revisiting Context Aggregation for Image Matting

ICML 2024poster

Traditional studies emphasize the significance of context information in improving matting performance. Consequently, deep learning-based matting methods delve into designing pooling or affinity-based context aggregation modules to achieve superior results. However, these modules cannot well handle…

2024

SpectralNeRF: Physically Based Spectral Rendering with Neural Radiance Field

AAAI 2024technical

In this paper, we propose SpectralNeRF, an end-to-end Neural Radiance Field (NeRF)-based architecture for high-quality physically based rendering from a novel spectral perspective. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning…

2023

CaPhy: Capturing Physical Properties for Animatable Human Avatars

ICCV 2023poster

We present CaPhy, a novel method for reconstructing animatable human avatars with realistic dynamic properties for clothing. Specifically, we aim for capturing the geometric and physical properties of the clothing from real observations. This allows us to apply novel poses to the human avatar with p…

Cited by 15PDFScholar
2023

Interactive Object Placement with Reinforcement Learning

ICML 2023poster

Object placement aims to insert a foreground object into a background image with a suitable location and size to create a natural composition. To predict a diverse distribution of placements, existing methods usually establish a one-to-one mapping from random vectors to the placements. However, thes…

Cited by 6SourcePDFScholar
2021

Distractor-Aware Fast Tracking via Dynamic Convolutions and MOT Philosophy

CVPR 2021poster

A practical long-term tracker typically contains three key properties, i.e., an efficient model design, an effective global re-detection strategy and a robust distractor awareness mechanism. However, most state-of-the-art long-term trackers (e.g., Pseudo and re-detecting based ones) do not take all…

Cited by 53PDFcodeScholar
2021

Efficient Regional Memory Network for Video Object Segmentation

CVPR 2021poster

Recently, several Space-Time Memory based networks have shown that the object cues (e.g. video frames as well as the segmented object masks) from the past frames are useful for segmenting objects in the current frame. However, these methods exploit the information from the memory by global-to-global…

Cited by 188PDFcodeScholar
2020

GRNet: Gridding Residual Network for Dense Point Cloud Completion

ECCV 2020poster

Estimating the complete 3D point cloud from an incomplete one is a key problem in many vision and robotics applications. Mainstream methods (e.g., PCN and TopNet) use Multi-layer Perceptrons (MLPs) to directly process point clouds, which may cause the loss of details because the structural and conte…

2020

Object-and-Action Aware Model for Visual Language Navigation

ECCV 2020poster

Vision-and-Language Navigation (VLN) is unique in that it requires turning relatively general natural-language instructions into robot agent actions, on the basis of visible environments. This requires to extract value from two very different types of natural-language information. The first is objec…

Cited by 133SourcePDFScholar
2020

Siamese Box Adaptive Network for Visual Tracking

CVPR 2020poster

Most of the existing trackers usually rely on either a multi-scale searching scheme or pre-defined anchor boxes to accurately estimate the scale and aspect ratio of a target. Unfortunately, they typically call for tedious and heuristic configurations. To address this issue, we propose a simple yet e…

Cited by 1031PDFcodeScholar
2019

Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View Images

ICCV 2019poster

Recovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input i…

Cited by 471PDFcodeScholar
2018

An Efficient Deep Convolutional Laplacian Pyramid Architecture for Cs Reconstruction At Low Sampling Ratios

ICASSP 2018accepted

The compressed sensing (CS) has been successfully applied to image compression in the past few years as most image signals are sparse in a certain domain. Several CS reconstruction models have been proposed and obtained superior performance. However, these methods suffer from blocking artifacts or r…

Cited by 0SourceScholar