← Search

Zhigang Wang

46 accepted papers

2026

Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose Estimation

AAAI 2026technical

Video-based human pose estimation has vast applications such as action recognition, sports analytics, and crime detection. However, this task is challenging as it involves interpreting both spatial context and temporal dynamics to accurately localize human anatomical keypoints in video sequences. Cu

Cited by 0SourcePDFScholar
2026

Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy

ICRA 2026poster

Diffusion-based policies have achieved remarkable results in robotic manipulation but often struggle to adapt rapidly in dynamic scenarios, leading to delayed responses or task failures. We present DCDP, a Dynamic Closed-Loop Diffusion Policy framework that integrates chunk-based action generation w…

2026

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation (VLN) requires the agent to navigate based on natural instructions. This task is challenging due to partial observability, which makes it difficult to align perception with language. Recent methods mitigate this by imagining future scenes, yet they rely on vision-based

Cited by 0SourcecodeScholar
2026

DiffusionPose: Markov-Optimized Diffusion Model for Human Pose Estimation

AAAI 2026technical

Video-based human pose estimation has long been a nontrivial task due to its dynamic nature and challenging detection scenarios such as occlusion and defocus. Inspired by the success of diffusion models, researchers have applied them to video pose estimation, outperforming traditional joint detectio

Cited by 0SourcePDFScholar
2026

Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in Videos

AAAI 2026technical

Video-based human pose estimation aims to localize keypoints across frames, enabling robust analysis of human motion in applications such as sports, surveillance, and healthcare. However, existing methods rely solely on visual cues, limiting their robustness in complex scenes involving occlusion, mo

Cited by 0SourcePDFScholar
2026

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models

AAAI 2026technical

Recent advancements have shown that the Mixture of Experts (MoE) approach significantly enhances the capacity of large language models (LLMs) and improves performance on downstream tasks. Building on these promising results, multi-modal large language models (MLLMs) have increasingly adopted MoE tec

Cited by 0SourcePDFScholar
2026

OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATION

ICLR 2026poster

Aerial Vision-Language Navigation (VLN) seeks to guide UAVs by leveraging language instructions and visual cues, establishing a new paradigm for human-UAV interaction. However, the collection of VLN data demands extensive human effort to construct trajectories and corresponding instructions, hinderi…

Cited by 0SourcecodeScholar
2025

AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations

ICCV 2025poster

Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditional VG, AerialVG poses new challenges, e.g., appearance-based grounding is insu…

2025

COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models

ICRA 2025

Leveraging the powerful reasoning capabilities of large language models (LLMs), recent LLM-based robot task planning methods yield promising results. However, they mainly focus on single or multiple homogeneous robots on simple tasks. Practically, complex long-horizon tasks always require collaborat

Cited by 41SourcecodeScholar
2025

Causal-Inspired Multitask Learning for Video-Based Human Pose Estimation

AAAI 2025technical

Video-based human pose estimation has long been a fundamental yet challenging problem in computer vision. Previous studies focus on spatio-temporal modeling through the enhancement of architecture design and optimization strategies. However, they overlook the causal relationships in the joints, lead…

Cited by 1SourcePDFScholar
2025

Cocube: a Tabletop Modular Multi-Robot Platform for Education and Research

ICRA 2025

This paper presents CoCube a tabletop modular robotics platform designed for robotics education and multirobot algorithm research. CoCube is characterized by its low cost low floors high ceilings and wide walls offering flexibility and broad applicability across various use cases. The platform compr

Cited by 0SourceScholar
2025

Efficient Diffusion as Low Light Enhancer

CVPR 2025poster

The computational burden of the iterative sampling process remains a major challenge in diffusion-based Low-Light Image Enhancement (LLIE). Current acceleration methods, whether training-based or training-free, often lead to significant performance degradation, highlighting the trade-off between per…

Cited by 0SourcePDFScholar
2025

FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset

CoRL 2025poster

Real-world manipulation datasets for robotic arms remain scarce due to the high costs, rigid hardware dependencies, and complex setup procedures associated with existing data collection methods. We introduce, a redesigned Universal Manipulation Interface (UMI) that addresses these challenges, enabli…

Cited by 0SourceScholar
2025

Generalize Audio Deepfake Algorithm Recognition via Attribution Enhancement

ICASSP 2025accepted

The development of voice cloning techniques has made forgery audios indistinguishable, posing an urgency to trace their sources. Many existing works focus on improving identification accuracy for audio deepfake algorithm recognition. However, most methods ignore the impact of complex information in…

Cited by 0SourceScholar
2025

Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding

AAAI 2025technical

3D Object Affordance Grounding aims to predict the functional regions on a 3D object and has laid the foundation for a wide range of applications in robotics. Recent advances tackle this problem via learning a mapping between 3D regions and a single human-object interaction image. However, the geome…

2025

MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation

ICCV 2025poster

In mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the nece…

Cited by 0SourcePDFScholar
2025

Multi-Grained Feature Pruning for Video-Based Human Pose Estimation

ICASSP 2025accepted

Human pose estimation, with its broad applications in action recognition and motion capture, has experienced significant advancements. However, current Transformer-based methods for video pose estimation often face challenges in managing redundant temporal information and achieving fine-grained perc…

Cited by 0SourceScholar
2025

Optimizing Human Pose Estimation Through Focused Human and Joint Regions

AAAI 2025technical

Human pose estimation has given rise to a broad spectrum of novel and compelling applications, including action recognition, sports analysis, as well as surveillance. However, accurate video pose estimation remains an open challenge. One aspect that has been overlooked so far is that existing method…

Cited by 1SourcePDFScholar
2025

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models

RSS 2025poster

In this paper, we claim that spatial understanding is the keypoint in robot manipulation, and propose SpatialVLA to explore effective spatial representations for the robot foundation model. Specifically, we propose Ego3D Position Encoding to inject 3D information into VLA’s input observations, and i…

Cited by 18PDFScholar
2025

SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos

AAAI 2025technical

Human pose estimation in videos remains a challenge, largely due to the reliance on extensive manual annotation of large datasets, which is expensive and labor-intensive. Furthermore, existing approaches often struggle to capture long-range temporal dependencies and overlook the complementary relati…

Cited by 1SourcePDFScholar
2025

Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

CVPR 2025poster

Learning a generalist robot that can effectively leverage prior knowledge for continuous skill acquisition remains significantly challenging. Despite the success of experience replay and parameter-efficient methods in maintaining knowledge across skills, naively applying these methods causes a failu…

Cited by 0SourcePDFScholar
2024

Any2Point: Empowering Any-modality Transformers for Efficient 3D Understanding

ECCV 2024poster

"Large foundation models have recently emerged as a prominent focus of interest, attaining superior performance in widespread scenarios. Due to the scarcity of 3D data, many efforts have been made to adapt pre-trained transformers from vision to 3D domains. However, such 2D-to-3D approaches are stil…

2024

Color Event Enhanced Single-Exposure HDR Imaging

AAAI 2024technical

Single-exposure high dynamic range (HDR) imaging aims to reconstruct the wide-range intensities of a scene by using its single low dynamic range (LDR) image, thus providing significant efficiency. Existing methods pay high attention to restoring the luminance by inversing the tone-mapping process, w…

2024

Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injection

IROS 2024poster

3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their effectiveness in fine-grained robotic manipulation tasks. To address…

Cited by 2SourcecodeScholar
2024

GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting

CVPR 2024highlight

In this paper we introduce GS-SLAM that first utilizes 3D Gaussian representation in the Simultaneous Localization and Mapping (SLAM) system. It facilitates a better balance between efficiency and accuracy. Compared to recent SLAM methods employing neural implicit representations our method utilizes…

2024

HPL-ESS: Hybrid Pseudo-Labeling for Unsupervised Event-based Semantic Segmentation

CVPR 2024poster

Event-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions which cannot be addressed by conventional RGB cameras. Since it is hard to annotate event data previous approaches rely on event-to-image recon…

Cited by 5SourcePDFScholar
2024

KOI: Accelerating Online Imitation Learning via Hybrid Key-state Guidance

CoRL 2024poster

Online Imitation Learning methods struggle with the gap between extensive online exploration space and limited expert trajectories, which hinder efficient exploration due to inaccurate task-aware reward estimation. Inspired by the findings from cognitive neuroscience that task decomposition coul…

Cited by 0SourceScholar
2024

Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMs

ICRA 2024poster

Generalizable articulated object manipulation is essential for home-assistant robots. Recent efforts focus on imitation learning from demonstrations or reinforcement learning in simulation, however, due to the prohibitive costs of real-world data collection and precise object simulation, it still re…

Cited by 23SourcecodeScholar
2024

LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Control and Rendering

NeurIPS 2024poster

This paper scales object-level reconstruction to complex scenes, advancing interactive scene reconstruction. We introduce two datasets, OmniSim and InterReal, featuring 28 scenes with multiple interactive objects. To tackle the challenge of inaccurate interactive motion recovery in complex scenes, w…

2024

Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models

AAAI 2024technical

The popularity of pre-trained large models has revolutionized downstream tasks across diverse fields, such as language, vision, and multi-modality. To minimize the adaption cost for downstream tasks, many Parameter-Efficient Fine-Tuning (PEFT) techniques are proposed for language and 2D image pre-tr…

2024

Revolutionizing Battery Disassembly: The Design and Implementation of a Battery Disassembly Autonomous Mobile Manipulator Robot(BEAM-1)

IROS 2024poster

The efficient disassembly of end-of-life electric vehicle batteries(EOL-EVBs) is crucial for green manufacturing and sustainable development. The current pre-programmed disassembly conducted by the Autonomous Mobile Manipulator Robot(AMMR) struggles to meet the disassembly requirements in dynamic en…

Cited by 5SourceScholar
2024

SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation

ICML 2024poster

Acquiring a multi-task imitation policy in 3D manipulation poses challenges in terms of scene understanding and action prediction. Current methods employ both 3D representation and multi-view 2D representation to predict the poses of the robot’s end-effector. However, they still require a considerab…

Cited by 11SourcePDFScholar
2024

X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge Transfer

AAAI 2024technical

The field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning tempo…

2023

Affordance-Driven Next-Best-View Planning for Robotic Grasping

CoRL 2023poster

Grasping occluded objects in cluttered environments is an essential component in complex robotic manipulation tasks. In this paper, we introduce an AffordanCE-driven Next-Best-View planning policy (ACE-NBV) that tries to find a feasible grasp for target object via continuously observing scenes from…

Cited by 14SourceScholar
2023

Fully Self-Supervised Depth Estimation From Defocus Clue

CVPR 2023poster

Depth-from-defocus (DFD), modeling the relationship between depth and defocus pattern in images, has demonstrated promising performance in depth estimation. Recently, several self-supervised works try to overcome the difficulties in acquiring accurate depth ground-truth. However, they depend on the…

2023

One-Shot High-Fidelity Talking-Head Synthesis With Deformable Neural Radiance Field

CVPR 2023poster

Talking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encou…

Cited by 54SourcePDFScholar
2023

Propagate and Calibrate: Real-Time Passive Non-Line-of-Sight Tracking

CVPR 2023poster

Non-line-of-sight (NLOS) tracking has drawn increasing attention in recent years, due to its ability to detect object motion out of sight. Most previous works on NLOS tracking rely on active illumination, e.g., laser, and suffer from high cost and elaborate experimental conditions. Besides, these te…

2023

Towards Nonlinear-Motion-Aware and Occlusion-Robust Rolling Shutter Correction

ICCV 2023poster

This paper addresses the problem of rolling shutter correction in complex nonlinear and dynamic scenes with extreme occlusion. Existing methods suffer from two main drawbacks. Firstly, they face challenges in estimating the accurate correction field due to the uniform velocity assumption, leading t…

Cited by 8PDFcodeScholar
2023

ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding

ICCV 2023poster

Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality and fail to weigh the relative importance of different views. In this paper, we propos…

Cited by 64PDFScholar
2022

Implicit Sample Extension for Unsupervised Person Re-Identification

CVPR 2022poster

Most existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters su…

Cited by 135PDFcodeScholar
2022

Self-Guided Hard Negative Generation for Unsupervised Person Re-Identification

IJCAI 2022poster

Recent unsupervised person re-identification (reID) methods mostly apply pseudo labels from clustering algorithms as supervision signals. Despite great success, this fashion is very likely to aggregate different identities with similar appearances into the same cluster. In result, the hard negative…

Cited by 12SourcePDFScholar
2021

Unsupervised Multi-Source Domain Adaptation for Person Re-Identification

CVPR 2021poster

Unsupervised domain adaptation (UDA) methods for person re-identification (re-ID) aim at transferring re-ID knowledge from labeled source data to unlabeled target data. Among these methods, the pseudo-label-based branch has achieved great success, whereas most of them only use limited data from a si…

Cited by 112PDFScholar
2020

Are We Ready for Service Robots? The OpenLORIS-Scene Datasets for Lifelong SLAM

ICRA 2020poster

Service robots should be able to operate autonomously in dynamic and daily changing environments over an extended period of time. While Simultaneous Localization And Mapping (SLAM) is one of the most fundamental problems for robotic autonomy, most existing SLAM works are evaluated with data sequence…

Cited by 174SourcecodeScholar
2020

Robotic Deep Rolling With Iterative Learning Motion and Force Control

RA-L 2020

Large industrial robots offer an attractive option for deep rolling in terms of cost and flexibility. These robots are typically designed for fast and precise motion, but may be commanded to perform force control by adjusting the position setpoint based on the measurements from a wrist-mounted force

Cited by 32SourceScholar