← Search

Zhenhua Wang

17 accepted papers

2025

VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware Tree

NeurIPS 2025poster

Video anomaly detection (VAD) focuses on identifying anomalies in videos. Su- pervised methods demand substantial in-domain training data and fail to deliver clear explanations for anomalies. In contrast, training-free methods leverage the knowledge reserves and language interactivity of large pre-t…

Cited by 0SourcecodeScholar
2023

CTVIS: Consistent Training for Online Video Instance Segmentation

ICCV 2023poster

The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/nega…

Cited by 46PDFcodeScholar
2022

3D Object Aided Self-Supervised Monocular Depth Estimation

IROS 2022poster

Monocular depth estimation has been actively studied in fields such as robot vision, autonomous driving, and 3D scene understanding. Given a sequence of color images, unsupervised learning methods based on the framework of Structure-From-Motion (SfM) simultaneously predict depth and camera relative…

Cited by 1SourceScholar
2022

CPGNet: Cascade Point-Grid Fusion Network for Real-Time LiDAR Semantic Segmentation

ICRA 2022poster

LiDAR semantic segmentation essential for advanced autonomous driving is required to be accurate, fast, and easy-deployed on mobile platforms. Previous point-based or sparse voxel-based methods are far away from real-time applications since time-consuming neighbor searching or sparse 3D convolution…

Cited by 34SourcecodeScholar
2022

ISDA: Position-Aware Instance Segmentation with Deformable Attention

ICASSP 2022accepted

Most instance segmentation models are not end-to-end trainable due to either the incorporation of proposal estimation (RPN) as a pre-processing or non-maximum suppression (NMS) as a post-processing. Here we propose a novel end-to-end instance segmentation method termed ISDA. It reshapes the task int…

Cited by 0SourceScholar
2022

PRNet: Point-Range Fusion Network for Real-Time LiDAR Semantic Segmentation

IJCAI 2022poster

Accurate and real-time LiDAR semantic segmentation is necessary for advanced autonomous driving systems. To guarantee a fast inference speed, previous methods utilize the highly optimized 2D convolutions to extract features on the range view (RV), which is the most compact representation of the LiDA…

Cited by 2SourcePDFScholar
2022

Robust and Accurate Multi-Agent SLAM with Efficient Communication for Smart Mobiles

ICRA 2022poster

In a long-term large-scenario application, the multi-agent collaborative SLAM is expected to improve the robustness and efficiency of executing tasks for mobile agents. In this paper, a multi-agent collaborative visual-inertial SLAM system is proposed based on a centralized client-server (CS) archit…

Cited by 8SourceScholar
2022

Sequential Multi-View Fusion Network for Fast LiDAR Point Motion Estimation

ECCV 2022poster

"The LiDAR point motion estimation, including motion state prediction and velocity estimation, is crucial for understanding a dynamic scene in autonomous driving. Recent 2D projection-based methods run in real-time by applying the well-optimized 2D convolution networks on either the bird’s-eye view…

Cited by 3SourcePDFScholar
2021

Consistency-Aware Graph Network for Human Interaction Understanding

ICCV 2021poster

Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main cause is that recent approaches learn human interactive relations via shallow graphical models…

Cited by 12PDFcodeScholar
2020

CalibRCNN: Calibrating Camera and LiDAR by Recurrent Convolutional Neural Network and Geometric Constraints

IROS 2020poster

In this paper, we present Calibration Recurrent Convolutional Neural Network (CalibRCNN) to infer a 6 degrees of freedom (DOF) rigid body transformation between 3D LiDAR and 2D camera. Different from the existing methods, our 3D-2D CalibRCNN not only uses the LSTM network to extract the temporal fea…

Cited by 72SourceScholar
2020

SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking

CVPR 2020oral

By decomposing the visual tracking task into two subproblems as classification for pixel category and regression for object bounding box at this pixel, we propose a novel fully convolutional Siamese network to solve visual tracking end-to-end in a per-pixel manner. The proposed framework SiamCAR con…

Cited by 980PDFcodeScholar
2020

SpSequenceNet: Semantic Segmentation Network on 4D Point Clouds

CVPR 2020poster

Point clouds are useful in many applications like autonomous driving and robotics as they provide natural 3D information of the surrounding environments. While there are extensive research on 3D point clouds, scene understanding on 4D point clouds, a series of consecutive 3D point clouds frames, is…

Cited by 129PDFScholar
2019

New Convex Relaxations for MRF Inference With Unknown Graphs

ICCV 2019poster

Treating graph structures of Markov random fields as unknown and estimating them jointly with labels have been shown to be useful for modeling human activity recognition and other related tasks. We propose two novel relaxations for solving this problem. The first is a linear programming (LP) relaxat…

Cited by 6PDFScholar
2018

A Hierarchical Model for Action Recognition Based on Body Parts

ICRA 2018poster

As increasing attention is paid on human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in human actions. Considering human actions as simultaneous motions of different body-parts…

Cited by 5SourceScholar
2017

Laplace gradient based Discriminative and Contrast Invertible descriptor

ICASSP 2017accepted

The performance of local descriptors such as SIFT drops under severe illumination changes. In this paper, we propose a Discriminative and Contrast Invertible (DCI) local feature descriptor. In order to increase the discriminative ability of the descriptor under illumination changes, a Laplace gradie…

Cited by 0SourceScholar