← Search

Wentao Zhu

29 accepted papers

2026

Electromagnetic Inverse Scattering from a Single Transmitter

CVPR 2026

Electromagnetic Inverse Scattering Problems (EISP) seek to reconstruct relative permittivity from scattered fields and are fundamental to applications like medical imaging. This inverse process is inherently ill-posed and highly nonlinear, making it particularly challenging, especially under sparse

Cited by 0SourcecodeScholar
2026

UNO! UNified Offline Training Paradigm for Learning Path Recommendation

AAAI 2026technical

With the wide adoption of online education platforms, adaptive learning systems have become increasingly important. Learning Path Recommendation (LPR) aims to dynamically adjust learning content to optimize learning efficiency based on individual student needs. However, current LPR methods suffer fr

Cited by 0SourcePDFScholar
2025

Aligning Human Motion Generation with Human Perceptions

ICLR 2025poster

Human motion generation is a critical task with a wide spectrum of applications. Achieving high realism in generated motions requires naturalness, smoothness, and plausibility. However, current evaluation metrics often rely on simple heuristics or distribution distances and do not align well with hu…

2025

Dynamic Model-Bank Test-Time Adaptation for Automatic Speech Recognition

EMNLP 2025

End-to-end automatic speech recognition (ASR) based on deep learning has achieved impressive progress in recent years. However, the performance of ASR foundation model often degrades significantly on out-of-domain data due to real-world domain shifts. Test-Time Adaptation (TTA) methods aim to mitiga

Cited by 0SourcePDFScholar
2025

Embodied Representation Alignment with Mirror Neurons

ICCV 2025poster

Mirror neurons are a class of neurons that activate both when an individual observes an action and when they perform the same action. This mechanism reveals a fundamental interplay between action understanding and embodied execution, suggesting that these two abilities are inherently connected. None…

Cited by 0SourcePDFScholar
2025

FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling

CVPR 2025highlight

Achieving realistic animated human avatars requires accurate modeling of pose-dependent clothing deformations. Existing learning-based methods heavily rely on the Linear Blend Skinning (LBS) of minimally-clothed human models like SMPL to model deformation. However, they struggle to handle loose clot…

Cited by 0SourcePDFScholar
2024

ScoreHypo: Probabilistic Human Mesh Estimation with Hypothesis Scoring

CVPR 2024poster

Monocular 3D human mesh estimation is an ill-posed problem characterized by inherent ambiguity and occlusion. While recent probabilistic methods propose generating multiple solutions little attention is paid to obtaining high-quality estimates from them. To address this limitation we introduce Score…

2023

3D Human Mesh Estimation From Virtual Markers

CVPR 2023poster

Inspired by the success of volumetric 3D pose estimation, some recent human mesh estimators propose to estimate 3D skeletons as intermediate representations, from which, the dense 3D meshes are regressed by exploiting the mesh topology. However, body shape information is lost in extracting skeletons…

2023

ChimpACT: A Longitudinal Dataset for Understanding Chimpanzee Behaviors

NeurIPS 2023poster

Understanding the behavior of non-human primates is crucial for improving animal welfare, modeling social behavior, and gaining insights into distinctively human and phylogenetically shared behaviors. However, the lack of datasets on non-human primate behavior hinders in-depth exploration of primate…

2023

Dichotomous Image Segmentation with Frequency Priors

IJCAI 2023poster

Dichotomous image segmentation (DIS) has a wide range of real-world applications and gained increasing research attention in recent years. In this paper, we propose to tackle DIS with informative frequency priors. Our model, called FP-DIS, stems from the fact that prior knowledge in the frequency do…

2023

Dynamic Inference With Grounding Based Vision and Language Models

CVPR 2023poster

Transformers have been recently utilized for vision and language tasks successfully. For example, recent image and language models with more than 200M parameters have been proposed to learn visual grounding in the pre-training step and show impressive results on downstream vision and language tasks.…

2023

GFPose: Learning 3D Human Pose Prior With Gradient Fields

CVPR 2023poster

Learning 3D human pose prior is essential to human-centered AI. Here, we present GFPose, a versatile framework to model plausible 3D human poses for various applications. At the core of GFPose is a time-dependent score network, which estimates the gradient on each body joint and progressively denois…

2023

MotionBERT: A Unified Perspective on Learning Human Motion Representations

ICCV 2023poster

We present a unified perspective on tackling various human-centric video tasks by learning human motion representations from large-scale and heterogeneous data resources. Specifically, we propose a pretraining stage in which a motion encoder is trained to recover the underlying 3D motion from noisy…

Cited by 215PDFcodeScholar
2023

Selective Structured State-Spaces for Long-Form Video Understanding

CVPR 2023poster

Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence (S4) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all image-tokens equall…

Cited by 126SourcePDFScholar
2023

Social Motion Prediction with Cognitive Hierarchies

NeurIPS 2023poster

Humans exhibit a remarkable capacity for anticipating the actions of others and planning their own actions accordingly. In this study, we strive to replicate this ability by addressing the social motion prediction problem. We introduce a new benchmark, a novel formulation, and a cognition-inspired f…

Cited by 9SourcePDFScholar
2022

Anti-Retroactive Interference for Lifelong Learning

ECCV 2022poster

"Humans can continuously learn new knowledge. However, machine learning models suffer from drastic dropping in performance on previous tasks after learning new tasks. Cognitive science points out that the competition of similar knowledge is an important cause of forgetting. In this paper, we design…

2022

CelebV-HQ: A Large-Scale Video Facial Attributes Dataset

ECCV 2022poster

"Large-scale datasets played an indispensable role in the recent success of face generation/editing and significantly facilitate the advances of emerging research fields. However, the academic community still lacks a video dataset with diverse facial attribute annotations, which is crucial for face-…

2022

Faster VoxelPose: Real-Time 3D Human Pose Estimation by Orthographic Projection

ECCV 2022poster

"While the voxel-based methods have achieved promising results for multi-person 3D pose estimation from multi-cameras, they suffer from heavy computation burdens, especially for large scenes. We present Faster VoxelPose to address the challenge by re-projecting the feature volume to the three two-di…

2022

MoCaNet: Motion Retargeting In-the-Wild via Canonicalization Networks

AAAI 2022technical

We present a novel framework that brings the 3D motion retargeting task from controlled environments to in-the-wild scenarios. In particular, our method is capable of retargeting body motion from a character in a 2D monocular video to a 3D character without using any motion capture system or 3D reco…

Cited by 15SourcePDFScholar
2022

Multi-Granularity Pruning for Model Acceleration on Mobile Devices

ECCV 2022poster

"For practical deep neural network design on mobile devices, it is essential to consider the constraints incurred by the computational resources and the inference latency in various applications. Among deep network acceleration approaches, pruning is a widely adopted practice to balance the computat…

Cited by 6SourcePDFScholar
2021

Shifted Chunk Transformer for Spatio-Temporal Representational Learning

NeurIPS 2021poster

Spatio-temporal representational learning has been widely adopted in various fields such as action recognition, video object segmentation, and action anticipation.Previous spatio-temporal representational learning approaches primarily employ ConvNets or sequential models, e.g., LSTM, to learn the in…

Cited by 43SourcePDFScholar
2021

Test-Time Training for Deformable Multi-Scale Image Registration

ICRA 2021poster

Registration is a fundamental task in medical robotics and is often a crucial step for many downstream tasks such as motion analysis, intra-operative tracking and image segmentation. Popular registration methods such as ANTs and NiftyReg optimize objective functions for each pair of images from scra…

Cited by 32SourceScholar
2020

Cycle-Consistent Adversarial Autoencoders for Unsupervised Text Style Transfer

COLING 2020main

Unsupervised text style transfer is full of challenges due to the lack of parallel data and difficulties in content preservation. In this paper, we propose a novel neural approach to unsupervised text style transfer which we refer to as Cycle-consistent Adversarial autoEncoders (CAE) trained from no…

Cited by 39SourcePDFScholar
2020

TransMoMo: Invariance-Driven Unsupervised Video Motion Retargeting

CVPR 2020poster

We present a lightweight video motion retargeting approach TransMoMo that is capable of transferring motion of a person in a source video realistically to another video of a target person. Without using any paired data for supervision, the proposed method can be trained in an unsupervised manner by…

Cited by 61PDFcodeScholar