← Search

Wayne Wu

52 accepted papers

2026

From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning

ICLR 2026poster

Navigation foundation models trained on massive web-scale data enable agents to generalize across diverse environments and embodiments. However, these models, which are trained solely on offline data, often lack the capacity to reason about the consequences of their actions or adapt through counterf…

Cited by 0SourceScholar
2026

Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion

ICRA 2026poster

Sidewalk micromobility is a promising solution for last-mile transportation, but current learning-based control methods struggle in complex urban environments. Imitation learning (IL) learns policies from human demonstrations, yet its reliance on fixed offline data often leads to compounding errors,…

2026

UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos

ICLR 2026poster

Urban embodied AI agents, ranging from delivery robots to quadrupeds, are increasingly populating our cities, navigating chaotic streets to provide last-mile connectivity. Training such agents requires diverse, high-fidelity urban environments to scale, yet existing human-crafted or procedurally gen…

Cited by 0SourceScholar
2025

Learning to Generate Diverse Pedestrian Movements from Web Videos with Noisy Labels

ICLR 2025poster

Understanding and modeling pedestrian movements in the real world is crucial for applications like motion forecasting and scene simulation. Many factors influence pedestrian movements, such as scene context, individual characteristics, and goals, which are often ignored by the existing human generat…

2025

MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention

CVPR 2025poster

Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling multiview attention to higher resolutions. In this paper, we…

2025

MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility

ICLR 2025spotlight

Public urban spaces such as streetscapes and plazas serve residents and accommodate social life in all its vibrant variations. Recent advances in robotics and embodied AI make public urban spaces no longer exclusive to humans. Food delivery bots and electric wheelchairs have started sharing sidewalk…

Cited by 1SourcePDFScholar
2025

Towards Autonomous Micromobility through Scalable Urban Simulation

CVPR 2025highlight

Micromobility, which utilizes lightweight devices moving in urban public spaces - such as delivery robots and electric wheelchairs - emerges as a promising alternative to vehicular mobility. Current micromobility depends mostly on human manual operation (in-person or remote control), which raises sa…

Cited by 1SourcePDFScholar
2025

Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation

CVPR 2025poster

Sim-to-real gap has long posed a significant challenge for robot learning in simulation, preventing the deployment of learned models in the real world. Previous work has primarily focused on domain randomization and system identification to mitigate this gap. However, these methods are often limited…

Cited by 4SourcePDFScholar
2024

CosmicMan: A Text-to-Image Foundation Model for Humans

CVPR 2024highlight

We present CosmicMan a text-to-image foundation model specialized for generating high-fidelity human images. Unlike current general-purpose foundation models that are stuck in the dilemma of inferior quality and text-image misalignment for humans CosmicMan enables generating photo-realistic human im…

2024

PaintHuman: Towards High-Fidelity Text-to-3D Human Texturing via Denoised Score Distillation

AAAI 2024technical

Recent advances in zero-shot text-to-3D human generation, which employ the human model prior (e.g., SMPL) or Score Distillation Sampling (SDS) with pre-trained text-to-image diffusion models, have been groundbreaking. However, SDS may provide inaccurate gradient directions under the weak diffusion g…

2024

Parameterization-driven Neural Surface Reconstruction for Object-oriented Editing in Neural Rendering

ECCV 2024poster

"The advancements in neural rendering have increased the need for techniques that enable intuitive editing of 3D objects represented as neural implicit surfaces. This paper introduces a novel neural algorithm for parameterizing neural implicit surfaces to simple parametric domains like spheres and p…

2023

CelebV-Text: A Large-Scale Facial Text-Video Dataset

CVPR 2023poster

Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents CelebV-Text, a large-scale, diverse, a…

2023

DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-Centric Rendering

ICCV 2023poster

Realistic human-centric rendering plays a key role in both computer vision and computer graphics. Rapid progress has been made in the algorithm aspect over the years, yet existing human-centric rendering datasets and benchmarks are rather impoverished in terms of diversity (e.g., outfit's fabric/mat…

Cited by 61PDFcodeScholar
2023

Filter-Recovery Network for Multi-Speaker Audio-Visual Speech Separation

ICLR 2023poster

In this paper, we systematically study the audio-visual speech separation task in a multi-speaker scenario. Given the facial information of each speaker, the goal of this task is to separate the corresponding speech from the mixed speech. The existing works are designed for speech separation in a co…

Cited by 5SourcePDFScholar
2023

MonoHuman: Animatable Human Neural Field From Monocular Video

CVPR 2023poster

Animating virtual avatars with free-view control is crucial for various applications like virtual reality and digital entertainment. Previous studies have attempted to utilize the representation power of the neural radiance field (NeRF) to reconstruct the human body from monocular videos. Recent wor…

Cited by 94SourcePDFScholar
2023

MotionBERT: A Unified Perspective on Learning Human Motion Representations

ICCV 2023poster

We present a unified perspective on tackling various human-centric video tasks by learning human motion representations from large-scale and heterogeneous data resources. Specifically, we propose a pretraining stage in which a motion encoder is trained to recover the underlying 3D motion from noisy…

Cited by 215PDFcodeScholar
2023

OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation

CVPR 2023poster

Recent advances in modeling 3D objects mostly rely on synthetic datasets due to the lack of large-scale real-scanned 3D databases. To facilitate the development of 3D perception, reconstruction, and generation in the real world, we propose OmniObject3D, a large vocabulary 3D object dataset with mass…

Cited by 214SourcePDFScholar
2023

RenderMe-360: A Large Digital Asset Library and Benchmarks Towards High-fidelity Head Avatars

NeurIPS 2023poster

Synthesizing high-fidelity head avatars is a central problem for computer vision and graphics. While head avatar synthesis algorithms have advanced rapidly, the best ones still face great obstacles in real-world scenarios. One of the vital causes is the inadequate datasets -- 1) current public data…

2023

SynBody: Synthetic Dataset with Layered Human Models for 3D Human Perception and Modeling

ICCV 2023poster

Synthetic data has emerged as a promising source for 3D human research as it offers low-cost access to large-scale human datasets. To advance the diversity and annotation quality of human models, we introduce a new synthetic dataset, SynBody, with three appealing features: 1) a clothed parametric hu…

Cited by 48PDFcodeScholar
2023

Text2Performer: Text-Driven Human Video Generation

ICCV 2023poster

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the appearance and motions of a target performer. Compared to general te…

Cited by 57PDFcodeScholar
2023

UnitedHuman: Harnessing Multi-Source Data for High-Resolution Human Generation

ICCV 2023poster

Human generation has achieved significant progress. Nonetheless, existing methods still struggle to synthesize specific regions such as faces and hands. We argue that the main reason is rooted in the training data. A holistic human dataset inevitably has insufficient and low-resolution information o…

Cited by 15PDFcodeScholar
2022

Audio-Driven Co-Speech Gesture Video Generation

NeurIPS 2022accept

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the image domain remains unsolved. In this work, we formally define and study this cha…

2022

CelebV-HQ: A Large-Scale Video Facial Attributes Dataset

ECCV 2022poster

"Large-scale datasets played an indispensable role in the recent success of face generation/editing and significantly facilitate the advances of emerging research fields. However, the academic community still lacks a video dataset with diverse facial attribute annotations, which is crucial for face-…

2022

Fast-Vid2Vid: Spatial-Temporal Compression for Video-to-Video Synthesis

ECCV 2022poster

"Video-to-Video synthesis (Vid2Vid) has achieved remarkable results on generating a photo-realistic video from a sequence of semantic maps. However, this pipeline suffers from high computational cost and long inference latency, which largely depends on two essential factors: 1) network architecture…

2022

Joint-Modal Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

ECCV 2022poster

"This paper focuses on the weakly-supervised audio-visual video parsing task, which aims to recognize all events belonging to each modality and localize their temporal boundaries. This task is challenging because only overall labels indicating the video events are provided for training. However, an…

2022

Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation

CVPR 2022poster

Generating speech-consistent body and gesture movements is a long-standing problem in virtual avatar creation. Previous studies often synthesize pose movement in a holistic manner, where poses of all joints are generated simultaneously. Such a straightforward pipeline fails to generate fine-grained…

Cited by 138PDFcodeScholar
2022

MoCaNet: Motion Retargeting In-the-Wild via Canonicalization Networks

AAAI 2022technical

We present a novel framework that brings the 3D motion retargeting task from controlled environments to in-the-wild scenarios. In particular, our method is capable of retargeting body motion from a character in a 2D monocular video to a 3D character without using any motion capture system or 3D reco…

Cited by 15SourcePDFScholar
2022

Progressive Attention on Multi-Level Dense Difference Maps for Generic Event Boundary Detection

CVPR 2022poster

Generic event boundary detection is an important yet challenging task in video understanding, which aims at detecting the moments where humans naturally perceive event boundaries. The main challenge of this task is perceiving various temporal variations of diverse event boundaries. To this end, this…

Cited by 20PDFcodeScholar
2022

Semantic-Aware Implicit Neural Audio-Driven Video Portrait Generation

ECCV 2022poster

"Animating high-fidelity video portrait with speech audio is crucial for virtual reality and digital entertainment. While most previous studies rely on accurate explicit structural information, recent works explore the implicit scene representation of Neural Radiance Fields (NeRF) for realistic gene…

2022

StyleGAN-Human: A Data-Centric Odyssey of Human Generation

ECCV 2022poster

"Unconditional human image generation is an important task in vision and graphics, enabling various applications in the creative industry. Existing studies in this field mainly focus on “network engineering” such as designing new components and objective functions. This work takes a data-centric per…

2022

TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial Editing

CVPR 2022poster

Recent advances like StyleGAN have promoted the growth of controllable facial editing. To address its core challenge of attribute decoupling in a single latent space, attempts have been made to adopt dual-space GAN for better disentanglement of style and content representations. Nonetheless, these m…

Cited by 73PDFcodeScholar
2021

Deceive D: Adaptive Pseudo Augmentation for GAN Training with Limited Data

NeurIPS 2021poster

Generative adversarial networks (GANs) typically require ample data for training in order to synthesize high-fidelity images. Recent studies have shown that training GANs with limited data remains formidable due to discriminator overfitting, the underlying cause that impedes the generator's converge…

2021

Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation

CVPR 2021poster

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information such as landmarks and 3D parameters, aiming to generate person…

Cited by 433PDFcodeScholar
2020

AOT: Appearance Optimal Transport Based Identity Swapping for Forgery Detection

NeurIPS 2020poster

Recent studies have shown that the performance of forgery detection can be improved with diverse and challenging Deepfakes datasets. However, due to the lack of Deepfakes datasets with large variance in appearance, which can be hardly produced by recent identity swapping methods, the detection algor…

2020

Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation

ECCV 2020poster

Depth information has proven to be a useful cue in the semantic segmentation of RGB-D images for providing a geometric counterpart to the RGB representation. Most existing works simply assume that depth measurements are accurate and well-aligned with the RGB pixels and models the problem as a cross-…

Cited by 431SourcePDFScholar
2020

DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection

CVPR 2020poster

We present our on-going effort of constructing a large- scale benchmark for face forgery detection. The first version of this benchmark, DeeperForensics-1.0, represents the largest face forgery detection dataset by far, with 60, 000 videos constituted by a total of 17.6 million frames, 10 times larg…

Cited by 586PDFcodeScholar
2020

TransMoMo: Invariance-Driven Unsupervised Video Motion Retargeting

CVPR 2020poster

We present a lightweight video motion retargeting approach TransMoMo that is capable of transferring motion of a person in a source video realistically to another video of a target person. Without using any paired data for supervision, the proposed method can be trained in an unsupervised manner by…

Cited by 61PDFcodeScholar
2019

Aggregation via Separation: Boosting Facial Landmark Detector With Semi-Supervised Style Translation

ICCV 2019poster

Facial landmark detection, or face alignment, is a fundamental task that has been extensively studied. In this paper, we investigate a new perspective of facial landmark detection and demonstrate it leads to further notable improvement. Given that any face images can be factored into space of style…

Cited by 102PDFcodeScholar
2019

FAB: A Robust Facial Landmark Detection Framework for Motion-Blurred Videos

ICCV 2019poster

Recently, facial landmark detection algorithms have achieved remarkable performance on static images. However, these algorithms are neither accurate nor stable in motion-blurred videos. The missing of structure information makes it difficult for state-of-the-art facial landmark detection algorithms…

Cited by 44PDFcodeScholar
2019

Make a Face: Towards Arbitrary High Fidelity Face Manipulation

ICCV 2019poster

Recent studies have shown remarkable success in face manipulation task with the advance of GANs and VAEs paradigms, but the outputs are sometimes limited to low-resolution and lack of diversity. In this work, we propose Additive Focal Variational Auto-encoder (AF-VAE), a novel approach that can arbi…

Cited by 86PDFScholar
2019

TransGaGa: Geometry-Aware Unsupervised Image-To-Image Translation

CVPR 2019poster

Unsupervised image-to-image translation aims at learning a mapping between two visual domains. However, learning a translation across large geometry variations al- ways ends up with failure. In this work, we present a novel disentangle-and-translate framework to tackle the complex objects image-to-i…

Cited by 135PDFScholar
2018

Look at Boundary: A Boundary-Aware Face Alignment Algorithm

CVPR 2018poster

We present a novel boundary-aware face alignment algorithm by utilising boundary lines as the geometric structure of a human face to help facial landmark localisation. Unlike the conventional heatmap based method and regression based method, our approach derives face landmarks from boundary lines wh…

2018

ReenactGAN: Learning to Reenact Faces via Boundary Transfer

ECCV 2018poster

We present a novel learning-based framework for face reenactment. The proposed method, known as ReenactGAN, is capable of transferring facial movements and expressions from an arbitrary person’s monocular video input to a target person’s video. Instead of performing a direct transfer in the pixel sp…