← Search

Wencheng Han

19 accepted papers

2026

DriveVLN: Towards Mapless Vision-and-Language Navigation in Autonomous Driving

CVPR 2026

Autonomous driving has made substantial progress recently, achieving reliable performance in most real-world environments. However, existing algorithms still depend heavily on high-definition maps, making them ineffective in mapless scenarios such as indoor parking lots. These limitations hinder sea

Cited by 0SourceScholar
2026

HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning

CVPR 2026

Recent advances in diffusion models have demonstrated impressive capability in generating high-quality images for simple prompts. However, when confronted with complex prompts involving multiple objects and hierarchical structures, existing models struggle to accurately follow instructions, leading

Cited by 0SourcecodeScholar
2026

Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion

AAAI 2026technical

Recent diffusion-based Single-image 3D portrait generation methods typically employ 2D diffusion models to provide multi-view knowledge, which is then distilled into 3D representations. However, these methods usually struggle to produce high-fidelity 3D models, frequently yielding excessively blurre

Cited by 0SourcePDFScholar
2025

ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions

ICCV 2025poster

3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting mechanism incorporating depth denoising. This enhances the robustnes…

2025

DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models

ICCV 2025poster

In this paper, we introduce DC (Decouple)-ControlNet, a highly flexible and precisely controllable framework for multi-condition image generation. The core idea behind DC-ControlNet is to decouple control conditions, transforming global control into a hierarchical system that integrates distinct ele…

Cited by 0SourcePDFScholar
2025

DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving

AAAI 2025technical

There are two crucial aspects of reliable autonomous driving systems: the reasoning behind decision-making and the precision of environmental perception. This paper introduces DME-Driver, a new autonomous driving system that enhances performance and robustness by fully leveraging the two crucial asp…

Cited by 25SourcePDFScholar
2025

Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution

CVPR 2025poster

Recovering high-quality depth maps from compressed sources has gained significant attention due to the limitations of consumer-grade depth cameras and the bandwidth restrictions during data transmission. However, current methods still suffer from two challenges. First, bit-depth compression produces…

2025

DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation

CVPR 2025poster

Autonomous driving evaluation requires simulation environments that closely replicate actual road conditions, including real-world sensory data and responsive feedback loops. However, many existing simulations need to predict waypoints along fixed routes on public datasets or synthetic photorealisti…

2025

Language Prompt for Autonomous Driving

AAAI 2025technical

A new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt. However, the progress of using language prompts in driving scenarios is stuck in a bottleneck due to the scarcity of paired prompt-instance data.…

2025

RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation

NeurIPS 2025poster

Synthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream perception tasks. We identify and quantify this critical issu…

Cited by 0SourceScholar
2025

Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation

ACL 2025finding

Text-to-image models are powerful for producing high-quality images based on given text prompts, but crafting these prompts often requires specialized vocabulary. To address this, existing methods train rewriting models with supervision from large amounts of manually annotated data and trained aesth…

Cited by 0SourcePDFScholar
2024

Leveraging Frame Affinity for sRGB-to-RAW Video De-rendering

CVPR 2024poster

Unprocessed RAW video has shown distinct advantages over sRGB video in video editing and computer vision tasks. However capturing RAW video is challenging due to limitations in bandwidth and storage. Various methods have been proposed to address similar issues in single image RAW capture through de-…

Cited by 2SourcePDFScholar
2024

RepVF: A Unified Vector Fields Representation for Multi-task 3D Perception

ECCV 2024poster

"Concurrent processing of multiple autonomous driving 3D perception tasks within the same spatiotemporal scene poses a significant challenge, in particular due to the computational inefficiencies and feature competition between tasks when using traditional multi-task learning approaches. This paper…

2023

Referring Multi-Object Tracking

CVPR 2023poster

Existing referring understanding tasks tend to involve the detection of a single text-referred object. In this paper, we propose a new and general referring understanding task, termed referring multi-object tracking (RMOT). Its core idea is to employ a language expression as a semantic cue to guide…

2023

Self-Supervised Monocular Depth Estimation by Direction-aware Cumulative Convolution Network

ICCV 2023poster

Monocular depth estimation is known as an ill-posed task that objects in a 2D image usually do not contain sufficient information to predict their depth. Thus, it acts differently from other tasks (e.g., classification and segmentation) in many ways. In this paper, we find that self-supervised monoc…

Cited by 25PDFcodeScholar
2023

Weakly Supervised Monocular 3D Object Detection Using Multi-View Projection and Direction Consistency

CVPR 2023poster

Monocular 3D object detection has become a mainstream approach in automatic driving for its easy application. A prominent advantage is that it does not need LiDAR point clouds during the inference. However, most current methods still rely on 3D point cloud data for labeling the ground truths used in…

2022

BRNet: Exploring Comprehensive Features for Monocular Depth Estimation

ECCV 2022poster

"Self-supervised monocular depth estimation has achieved promising performance recently. A consensus is that high-resolution inputs often yield better results. However, we find that the performance gap between high and low resolutions lies in the inappropriate feature representation of the widely us…

2021

Learning To Fuse Asymmetric Feature Maps in Siamese Trackers

CVPR 2021poster

Recently, Siamese-based trackers have achieved promising performance in visual tracking. Most recent Siamese-based trackers typically employ a depth-wise cross-correlation (DW-XCorr) to obtain multi-channel correlation information from the two feature maps (target and search region). However, DW-XCo…

Cited by 97PDFcodeScholar