← Search

Hanli Wang

12 accepted papers

2026

An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving

CVPR 2026

Panoptic occupancy prediction aims to jointly infer voxel-wise semantics and instance identities within a unified 3D scene representation. Nevertheless, progress in this field remains constrained by the absence of high-quality 3D mesh resources, instance-level annotations, and physically consistent

Cited by 1SourceScholar
2026

Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMs

CVPR 2026

Recently, multimodal large language models (MLLMs) have achieved remarkable success in general multimodal tasks. Increasing attention has been given to leveraging MLLMs for fine-grained visual understanding, such as region-level captioning and pixel-level grounding. However, most existing approaches

Cited by 0SourceScholar
2025

Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous Driving

AAAI 2025technical

Autonomous driving is a challenging task that requires perceiving and understanding the surrounding environment for safe trajectory planning. While existing vision-based end-to-end models have achieved promising results, these methods are still facing the challenges of vision understanding, decision…

2025

MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction

ICLR 2025poster

The construction of vectorized high-definition map typically requires capturing both category and geometry information of map elements. Current state-of-the-art methods often adopt solely either point-level or instance-level representation, overlooking the strong intrinsic relationship between point…

Cited by 3SourcePDFScholar
2024

ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance Field

AAAI 2024technical

Neural Radiance Fields (NeRF) have demonstrated impressive potential in synthesizing novel views from dense input, however, their effectiveness is challenged when dealing with sparse input. Existing approaches that incorporate additional depth or semantic supervision can alleviate this issue to an e…

2024

DDR: Exploiting Deep Degradation Response as Flexible Image Descriptor

NeurIPS 2024poster

Image deep features extracted by pre-trained networks are known to contain rich and informative representations. In this paper, we present Deep Degradation Response (DDR), a method to quantify changes in image deep features under varying degradation conditions. Specifically, our approach facilitates…

2024

Misalignment-Robust Frequency Distribution Loss for Image Transformation

CVPR 2024poster

This paper aims to address a common challenge in deep learning-based image transformation methods such as image enhancement and super-resolution which heavily rely on precisely aligned paired datasets with pixel-level alignments. However creating precisely aligned paired images presents significant…

2020

Enhanced Action Tubelet Detector for Spatio-Temporal Video Action Detection

ICASSP 2020accepted

Current spatio-temporal action detection methods usually employ a two-stream architecture, a RGB stream for raw images and an auxiliary motion stream for optical flow. Training is required individually for each stream and more efforts are necessary to improve the precision of RGB stream. To this end…

Cited by 0SourceScholar
2016

Blind image quality assessment for multiply distorted images via convolutional neural networks

ICASSP 2016accepted

The past decade has witnessed a growing development of Image Quality Assessment (IQA) techniques. However, the researches of IQA with multiple distortion types are still limited especially on blind image quality assessment methods. In this paper, a Convolutional Neural Network (CNN) based method is…

Cited by 0SourceScholar
2016

Real-Time Action Recognition With Enhanced Motion Vector CNNs

CVPR 2016poster

The deep two-stream architecture exhibited excellent performance on video based action recognition. The most computationally expensive step in this approach comes from the calculation of optical flow which prevents it to be real-time. This paper accelerates this architecture by replacing optical flo…

Cited by 546PDFcodeScholar