← Search

Haoyang Zhang

16 accepted papers

2026

Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving

CVPR 2026

3D semantic occupancy prediction is crucial for autonomous driving perception, offering comprehensive geometric scene understanding and semantic recognition. However, existing methods struggle with geometric misalignment in view transformation due to lack of pixel-level accurate depth estimation, an

Cited by 0SourceScholar
2026

MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs

CVPR 2026

3D hand pose estimation (HPE) from sparse inertial measurement units (IMUs) has shown great potential in human-computer interaction. However, due to the significant semantic gap between sparse local motion information and structured global pose information, estimating hand poses from sparse IMU sign

Cited by 0SourceScholar
2026

OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition

CVPR 2026

Online micro gesture recognition from hand skeletons is critical for VR/AR interaction but faces challenges due to limited public datasets and task-specific algorithms. Micro gestures involve subtle motion patterns, which make constructing datasets with precise skeletons and frame-level annotations

Cited by 0SourceScholar
2026

UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

CVPR 2026

Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between input images and rendered outputs, or multi-view consistency constraints, as the driving force to optimize networks and p

Cited by 0SourceScholar
2025

Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage

NeurIPS 2025poster

We present the design and implementation of a new lifetime-aware tensor offloading framework for GPU memory expansion using low-cost PCIe-based solid-state drives (SSDs). Our framework, TERAIO, is developed explicitly for large language model (LLM) training with multiple GPUs and multiple SSDs. Its…

Cited by 0SourceScholar
2025

Generating Multimodal Driving Scenes via Next-Scene Prediction

CVPR 2025poster

Generative models in Autonomous Driving (AD) enable diverse scenario creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multi…

2025

Hierarchical-aware Orthogonal Disentanglement Framework for Fine-grained Skeleton-based Action Recognition

ICCV 2025poster

In recent years, skeleton-based action recognition has gained significant attention due to its robustness in varying environmental conditions. However, most existing methods struggle to distinguish fine-grained actions due to subtle motion features, minimal inter-class variation, and they often fail…

Cited by 0SourcePDFScholar
2024

Leveraging the efficiency of multi-task robot manipulation via task-evoked planner and reinforcement learning

ICRA 2024poster

Multi-task learning has expanded the boundaries of robotic manipulation, enabling the execution of increasingly complex tasks. However, policies learned through reinforcement learning exhibit limited generalization and narrow distributions, which restrict their effectiveness in multi-task training.…

Cited by 0SourceScholar
2024

Online Trajectory Generation With Local Replanning for 7-DoF Serial Manipulator in Unforeseen Dynamic Environments

RA-L 2024

In this letter, we focus on online motion planning for manipulators in dynamic obstacle environments. An analytical geometry-based inverse kinematics solution for generalized types of 7-DoF anthropomorphic manipulators is presented to work as the basis of high-efficiency collision avoidance planning

Cited by 3SourceScholar
2024

Symphonize 3D Semantic Scene Completion with Contextual Instance Queries

CVPR 2024poster

3D Semantic Scene Completion (SSC) has emerged as a nascent and pivotal undertaking in autonomous driving aiming to predict the voxel occupancy within volumetric scenes. However prevailing methodologies primarily focus on voxel-wise feature aggregation while neglecting instance semantics and scene c…

2023

Towards Safe and Aggressive Motion Generation for Dynamic Targets Pick-and-Place

IROS 2023poster

In this paper, we present a framework to generate time-optimal trajectories for dynamic target pick-and-place tasks. We develop an optimization-based trajectory generation method for manipulators, which can conduct spatial-temporal deformation under user-defined requirements. We formulate the proble…

Cited by 2SourceScholar
2022

InsPro: Propagating Instance Query and Proposal for Online Video Instance Segmentation

NeurIPS 2022accept

Video instance segmentation (VIS) aims at segmenting and tracking objects in videos. Prior methods typically generate frame-level or clip-level object instances first and then associate them by either additional tracking heads or complex instance matching algorithms. This explicit instance associati…

Cited by 19SourcePDFScholar
2022

PanopticDepth: A Unified Framework for Depth-Aware Panoptic Segmentation

CVPR 2022poster

This paper presents a unified framework for depth-aware panoptic segmentation (DPS), which aims to reconstruct 3D scene with instance-level semantics from one single image. Prior works address this problem by simply adding a dense depth regression head to panoptic segmentation (PS) networks, resulti…

Cited by 29PDFcodeScholar
2021

Evaluating the Impact of Semantic Segmentation and Pose Estimation on Dense Semantic SLAM

IROS 2021poster

Recent Semantic SLAM methods combine classical geometry-based estimation with deep learning-based object detection or semantic segmentation. In this paper we evaluate the quality of semantic maps generated by state-of-the-art class-and instance-aware dense semantic SLAM algorithms whose codes are pu…

Cited by 11SourceScholar