← Search

Ailing Zeng

40 accepted papers

2026

Gloria: Consistent Character Video Generation via Content Anchors

CVPR 2026

Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non-character-centric information

Cited by 0SourceScholar
2026

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

ICLR 2026poster

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most notably video generation (ViGen), have demonstrated remarkable generalization in…

Cited by 0SourcecodeScholar
2025

DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior

ICCV 2025poster

We present DPoser-X, a diffusion-based prior model for 3D whole-body human poses. Building a versatile and robust full-body human pose prior remains challenging due to the inherent complexity of articulated human poses and the scarcity of high-quality whole-body pose datasets. To address these limit…

Cited by 0SourcePDFScholar
2025

IDOL: Instant Photorealistic 3D Human Creation from a Single Image

CVPR 2025poster

Creating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the…

2025

MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls

AAAI 2025technical

Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to process different condition modalities presents two main challenges: motion distribution drifts across di…

2025

SkillMimic: Learning Basketball Interaction Skills from Demonstrations

CVPR 2025highlight

Traditional reinforcement learning methods for human-object interaction (HOI) rely on labor-intensive, manually designed skill rewards that do not generalize well across different interactions. We introduce SkillMimic, a unified data-driven framework that fundamentally changes how agents learn inter…

2025

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

NeurIPS 2025poster

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality 1080P human speech videos w…

Cited by 0SourceScholar
2025

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation

ICML 2025poster

Although significant advancements have been achieved in the progress of keypoint-guided Text-to-Image diffusion models, existing mainstream keypoint-guided models encounter challenges in controlling the generation of more general non-rigid objects beyond humans (e.g., animals). Moreover, it is diffi…

Cited by 0SourcePDFScholar
2024

AddMe: Zero-shot Group-photo Synthesis by Inserting People into Scenes

ECCV 2024poster

"While large text-to-image diffusion models have made significant progress in high-quality image generation, challenges persist when users insert their portraits into existing photos, especially group photos. Concretely, existing customization methods struggle to insert facial identities at desired…

2024

AiOS: All-in-One-Stage Expressive Human Pose and Shape Estimation

CVPR 2024poster

Expressive human pose and shape estimation (a.k.a. 3D whole-body mesh recovery) involves the human body hand and expression estimation. Most existing methods have tackled this task in a two-stage manner first detecting the human body part with an off-the-shelf detection model and then inferring the…

2024

Bridging the Gap Between Human Motion and Action Semantics via Kinematics Phrases

ECCV 2024poster

"Motion understanding aims to establish a reliable mapping between motion and action semantics, while it is a challenging many-to-many problem. An abstract action semantic (i.e., walk forwards) could be conveyed by perceptually diverse motions (walking with arms up or swinging). In contrast, a motio…

2024

DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation

CVPR 2024poster

We propose DiffSHEG a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation. While previous works focused on co-speech gesture or expression generation individually the joint generation of synchronized expressions and gestures remains barely explored. To address th…

Cited by 37SourcePDFScholar
2024

FreeMan: Towards Benchmarking 3D Human Pose Estimation under Real-World Conditions

CVPR 2024poster

Estimating the 3D structure of the human body from nat- ural scenes is a fundamental aspect of visual perception. 3D human pose estimation is a vital step in advancing fields like AIGC and human-robot interaction serving as a crucial tech- nique for understanding and interacting with human actions i…

2024

GPAvatar: Generalizable and Precise Head Avatar from Image(s)

ICLR 2024poster

Head avatar reconstruction, crucial for applications in virtual reality, online meetings, gaming, and film industries, has garnered substantial attention within the computer vision community. The fundamental objective of this field is to faithfully recreate the head avatar and precisely control expr…

2024

HumanTOMATO: Text-aligned Whole-body Motion Generation

ICML 2024poster

This work targets a novel text-driven **whole-body** motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and body motions simultaneously. Previous works on text-driven motion generation…

2024

MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

NeurIPS 2024poster

Sora's high-motion intensity and long consistent videos have significantly impacted the field of video generation, attracting unprecedented attention. However, existing publicly available datasets are inadequate for generating Sora-like videos, as they mainly contain short videos with low motion int…

Cited by 42SourcePDFScholar
2024

Open-World Human-Object Interaction Detection via Multi-modal Prompts

CVPR 2024poster

In this paper we develop MP-HOI a powerful Multi-modal Prompt-based HOI detector designed to leverage both textual descriptions for open-set generalization and visual exemplars for handling high ambiguity in descriptions realizing HOI detection in the open world. Specifically it integrates visual pr…

2024

PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of Code

ICLR 2024poster

Text-guided diffusion models have revolutionized image generation and editing, offering exceptional realism and diversity. Specifically, in the context of diffusion-based editing, where a source image is edited according to a target prompt, the process commences by acquiring a noisy latent vector co…

Cited by 111SourcePDFScholar
2023

A Comprehensive Benchmark for Neural Human Radiance Fields

NeurIPS 2023poster

The past two years have witnessed a significant increase in interest concerning NeRF-based human body rendering. While this surge has propelled considerable advancements, it has also led to an influx of methods and datasets. This explosion complicates experimental settings and makes fair comparisons…

Cited by 2SourcePDFScholar
2023

Are Transformers Effective for Time Series Forecasting?

AAAI 2023technical

Recently, there has been a surge of Transformer-based solutions for the long-term time series forecasting (LTSF) task. Despite the growing performance over the past few years, we question the validity of this line of research in this work. Specifically, Transformers is arguably the most successful s…

2023

DreamWaltz: Make a Scene with Complex 3D Animatable Avatars

NeurIPS 2023poster

We present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects, creating high-quality and animatable 3D avatars remains chall…

2023

Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation

ICLR 2023poster

This paper presents a novel end-to-end framework with Explicit box Detection for multi-person Pose estimation, called ED-Pose, where it unifies the contextual learning between human-level (global) and keypoint-level (local) information. Different from previous one-stage methods, ED-Pose re-considers…

2023

From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft Labels

ICCV 2023poster

Knowledge Distillation (KD) uses the teacher's prediction logits as soft labels to guide the student, while self-KD does not need a real teacher to require the soft labels. This work unifies the formulations of the two tasks by decomposing and reorganizing the generic KD loss into a Normalized KD (N…

Cited by 111PDFcodeScholar
2023

Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes

CVPR 2023poster

Humans have long been recorded in a variety of forms since antiquity. For example, sculptures and paintings were the primary media for depicting human beings before the invention of cameras. However, most current human-centric computer vision tasks like human pose estimation and human image generati…

2023

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation

ICCV 2023oral

Controllable human image generation (HIG) has attracted significant attention from academia and industry for its numerous real-life applications. State-of-the-art solutions, such as ControlNet and T2I-Adapter, introduce an additional learnable branch on top of the frozen pre-trained stable diffusion…

Cited by 90PDFcodeScholar
2023

Lite DETR: An Interleaved Multi-Scale Encoder for Efficient DETR

CVPR 2023poster

Recent DEtection TRansformer-based (DETR) models have obtained remarkable performance. Its success cannot be achieved without the re-introduction of multi-scale feature fusion in the encoder. However, the excessively increased tokens in multi-scale features, especially for about 75% of low-level fea…

2023

One-Stage 3D Whole-Body Mesh Recovery With Component Aware Transformer

CVPR 2023poster

Whole-body mesh recovery aims to estimate the 3D human body, face, and hands parameters from a single image. It is challenging to perform this task with a single network due to resolution issues, i.e., the face and hands are usually located in extremely small regions. Existing works usually detect h…

2023

SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation

NeurIPS 2023poster

Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods still depend largely on a confined set of training datasets. In this work, we investigate scaling up EHPS towards…

2022

DeciWatch: A Simple Baseline for 10× Efficient 2D and 3D Pose Estimation

ECCV 2022poster

"This paper proposes a simple baseline framework for video-based 2D/3D human pose estimation that can achieve 10 times efficiency improvement over existing works without any performance degradation, named DeciWatch. Unlike current solutions that estimate each frame in a video, DeciWatch introduces a…

2022

HuMMan: Multi-modal 4D Human Dataset for Versatile Sensing and Modeling

ECCV 2022poster

"4D human sensing and modeling are fundamental tasks in vision and graphics with numerous applications. With the advances of new sensors and algorithms, there is an increasing demand for more versatile datasets. In this work, we contribute HuMMan, a large-scale multi-modal 4D human dataset with 1000…

Cited by 125SourcePDFScholar
2022

SCINet: Time Series Modeling and Forecasting with Sample Convolution and Interaction

NeurIPS 2022accept

One unique property of time series is that the temporal relations are largely preserved after downsampling into two sub-sequences. By taking advantage of this property, we propose a novel neural network architecture that conducts sample convolution and interaction for temporal modeling and forecasti…

2022

SmoothNet: A Plug-and-Play Network for Refining Human Poses in Videos

ECCV 2022poster

"When analyzing human motion videos, the output jitters from existing pose estimators are highly-unbalanced with varied estimation errors across frames. Most frames in a video are relatively easy to estimate and only suffer from slight jitters. In contrast, for rarely seen or occluded actions, the e…

2022

T-WaveNet: A Tree-Structured Wavelet Neural Network for Time Series Signal Analysis

ICLR 2022poster

Time series signal analysis plays an essential role in many applications, e.g., activity recognition and healthcare monitoring. Recently, features extracted with deep neural networks (DNNs) have shown to be more effective than conventional hand-crafted ones. However, most existing solutions rely sol…

Cited by 16SourcePDFScholar
2021

Human Pose Regression With Residual Log-Likelihood Estimation

ICCV 2021poster

Heatmap-based methods dominate in the field of human pose estimation by modelling the output distribution through likelihood heatmaps. In contrast, regression-based methods are more efficient but suffer from inferior performance. In this work, we explore maximum likelihood estimation (MLE) to develo…

Cited by 277PDFcodeScholar
2021

Information Bottleneck Approach to Spatial Attention Learning

IJCAI 2021poster

The selective visual attention mechanism in the human visual system (HVS) restricts the amount of information to reach visual awareness for perceiving natural scenes, allowing near real-time information processing with limited computational capacity. This kind of selectivity acts as an ‘Information…

2021

Learning Skeletal Graph Neural Networks for Hard 3D Pose Estimation

ICCV 2021poster

Various deep learning techniques have been proposed to solve the single-view 2D-to-3D pose estimation problem. While the average prediction accuracy has been improved significantly over the years, the performance on hard poses with depth ambiguity, self-occlusion, and complex or rare poses is still…

Cited by 163PDFScholar
2020

SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach

ECCV 2020poster

Human poses that are rare or unseen in a training set are challenging for a network to predict. Similar to the long-tailed distribution problem in visual recognition, the small number of examples for such poses limits the ability of networks to model them. Interestingly, local pose distributions suf…