← Search

Zhenzhi Wang

8 accepted papers

2026

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

ICLR 2026poster

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios that multiple concepts could…

Cited by 0SourceScholar
2025

Multi-identity Human Image Animation with Structural Video Diffusion

ICCV 2025poster

Generating human videos from a single image while ensuring high visual quality and precise control is a challenging task, especially in complex scenarios involving multiple individuals and interactions with objects. Existing methods, while effective for single-human cases, often fail to handle the i…

2024

HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation

NeurIPS 2024poster

Human image animation involves generating videos from a character photo, allowing user control and unlocking the potential for video and movie production. While recent approaches yield impressive results using high-quality training data, the inaccessibility of these datasets hampers fair and transpa…

2024

InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint

NeurIPS 2024poster

Text-conditioned motion synthesis has made remarkable progress with the emergence of diffusion models. However, the majority of these motion diffusion models are primarily designed for a single character and overlook multi-human interactions. In our approach, we strive to explore this problem by syn…

2023

MatrixCity: A Large-scale City Dataset for City-scale Neural Rendering and Beyond

ICCV 2023poster

Neural radiance fields (NeRF) and its subsequent variants have led to remarkable progress in neural rendering. While most of recent neural rendering works focus on objects and small-scale scenes, developing neural rendering methods for city-scale scenes is of great potential in many real-world appli…

Cited by 78PDFcodeScholar
2022

Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding

AAAI 2022technical

Temporal grounding aims to localize a video moment which is semantically aligned with a given natural language query. Existing methods typically apply a detection or regression pipeline on the fused representation with the research focus on designing complicated prediction heads or fusion strategies…

2021

MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions

ICCV 2021poster

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions. This paper aims to present a new multi-person dataset of spat…

Cited by 124PDFcodeScholar
2020

Boundary-Aware Cascade Networks for Temporal Action Segmentation

ECCV 2020poster

Identifying human action segments in an untrimmed video is still challenging due to boundary ambiguity and over-segmentation issues. To address these problems, we present a new boundary-aware cascade network by introducing two novel components. First, we devise a new cascading paradigm, called Stage…