← Search

Hao Gao

22 accepted papers

2026

Instance-level Visual Active Tracking with Occlusion-Aware Planning

CVPR 2026

Visual Active Tracking (VAT) aims to control cameras to follow a target in 3D space, which is critical for applications like drone navigation and security surveillance. However, it faces two key bottlenecks in real-world deployment: confusion from visually similar distractors caused by insufficient

Cited by 0SourcecodeScholar
2026

Point Cloud Quality Assessment via Multi-View Structure-Aware Feature Fusion

AAAI 2026technical

Point cloud quality assessment (PCQA) is essential for reliable 3D visual applications. While point-based methods face challenges in characterizing distortions due to point cloud disorder, projection-based approaches offer better efficiency but suffer from geometric distortion insensitivity and text

Cited by 0SourcePDFScholar
2026

VADv2: End-to-End Autonomous Driving via Probabilistic Planning

ICLR 2026poster

Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of planning make it challenging. Existing learning-based planning methods follow a deterministic paradigm to directly regress the action, failing to cope with t…

Cited by 0SourcecodeScholar
2025

Adaptive Skeleton Prompt Tuning for Cross-Dataset 3D Human Pose Estimation

ICASSP 2025accepted

Inconsistency of distributions in human actions and camera viewpoints can lead to significant deviations when the pre-trained 3D pose estimators are tested on cross-datasets. In practical applications, the estimators usually follow the standard full fine-tuning paradigm on the target dataset, which…

Cited by 0SourceScholar
2025

C2F-Planner: Interaction-Aware Coarse-to-Fine Planning for Autonomous Vehicles

RA-L 2025

Ensuring safe and socially compliant driving is essential for autonomous vehicle planning. However, one of the significant challenges remains the performance bottleneck caused by interaction uncertainty in complex traffic scenarios. Traditional planning algorithms typically account for all traffic p

Cited by 0SourcecodeScholar
2025

Diffusion Models are Good Unsupervised Class-agnostic Shape Part Segmentators

ICASSP 2025accepted

Shape part segmentation is a critical task in computer graphics and robotics. However, traditional supervised methods rely heavily on large amounts of labeled data, which poses significant challenges in many real-world scenarios where such data is often scarce or difficult to obtain. To address this…

Cited by 0SourceScholar
2025

Enhancing Autonomous Vehicle Planning With a Robust Fault-Tolerant Mechanism for Action-Induced Agent Detection

ICASSP 2025accepted

In autonomous driving, accurately identifying traffic participants that may influence vehicle behavior is crucial for effective system planning. To address this challenge, we propose a fault-tolerant mechanism for detecting action-induced objects, which significantly improves decision-making perform…

Cited by 0SourceScholar
2025

RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning

NeurIPS 2025poster

Existing end-to-end autonomous driving (AD) algorithms typically follow the Imitation Learning (IL) paradigm, which faces challenges such as causal confusion and an open-loop gap. In this work, we propose RAD, a 3DGS-based closed-loop Reinforcement Learning (RL) framework for end-to-end Autonomous D…

Cited by 0SourcecodeScholar
2025

RFEM: Remote Feature Enhancement Module for Target Detection

ICASSP 2025accepted

The research and development of dense crowd detection technology have always been one of the hot and challenging topics in the field of computer vision. DETR-like models have shown good performance in both training efficiency and inference capabilities. Nevertheless, as the optimization proceeds, th…

Cited by 0SourceScholar
2024

Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning

ICML 2024poster

In offline reinforcement learning, the challenge of out-of-distribution (OOD) is pronounced. To address this, existing methods often constrain the learned policy through policy regularization. However, these methods often suffer from the issue of unnecessary conservativeness, hampering policy improv…

2024

DeformMLP: Dynamic Large-Scale Receptive Field MLP Networks for Human Motion Prediction

ICASSP 2024accepted

Predicting human motion requires addressing dependencies and errors for pose forecasting from sequences. The transformer’s self-attention aids this, but its complexity poses computational challenges. We present an efficient DeformMLP network without self-attention, using fully connected layers. Defo…

Cited by 0SourceScholar
2024

Geometry Compression Artifact Removal for V-PCC over a Wide Bitrate Range

ICASSP 2024accepted

In video-based point cloud compression (V-PCC), point clouds are generated as videos via patch projection to be compressed using video coding techniques. However, a large number of filled empty pixels in the videos creates a fake context, which reduces the noise prediction accuracy in compression ar…

Cited by 0SourceScholar
2024

Local Optimization Networks for Multi-View Multi-Person Human Posture Estimation

ICASSP 2024accepted

With the growing applicability of multi-view multi-person 3D human pose estimation across diverse scenarios, the impact of external environmental factors and occlusion on accuracy has garnered substantial attention. In this research, we introduce a novel approach to multi-view multi-person 3D human…

Cited by 0SourceScholar
2023

Cross-Modal Optical Flow Estimation via Modality Compensation and Alignment

ICASSP 2023accepted

Cross-modal optical flow estimation aims to predict motion fields between two frames collected from different modalities, recently attracting intensive attention. However, a substantial yet challenging problem is how to match images across a large modal discrepancy. In this paper, we propose a modal…

Cited by 0SourceScholar
2023

Learning Hybrid Representations of Semantics and Distortion for Blind Image Quality Assessment

ICASSP 2023accepted

Recently, some studies have shown that semantic and distortion representations both benefit the evaluation of image quality. However, the images of existing synthetic distortion databases are annotated with subjective quality scores and distortion types, lacking labels with semantic objects. Therefo…

Cited by 0SourceScholar
2023

Learning Scene Flow from 3d Point Clouds with Cross-Transformer and Global Motion Cues

ICASSP 2023accepted

Scene flow estimation is critical for real-world vision problems such as autonomous driving and augmented reality. Due to the popularity of 3D LiDAR sensors, scene flow estimation from 3D point clouds arouses increasing attention. Existing methods usually use a flow embedding-based layer to find cor…

Cited by 0SourceScholar
2023

ψ-Net: Point Structural Information Network for No-Reference Point Cloud Quality Assessment

ICASSP 2023accepted

The human vision system is highly adapted to extract structural information from the viewed scenes. The irregularity of point clouds makes the extraction of structural information containing both color and geometry an important challenge for point cloud quality assessment (PCQA). This paper proposes…

Cited by 0SourceScholar