← Search

Yuhang Lu

14 accepted papers

2026

DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving

CVPR 2026

Multimodal Large Language Models (MLLMs) are rapidly becoming the intelligence brain of end-to-end autonomous driving systems. A key challenge is to assess whether MLLMs can truly understand and follow complex real-world traffic rules. However, existing benchmarks mainly focus on single-rule scenari

Cited by 0SourceScholar
2026

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-turn Dialogue

ICML 2026poster

Multimodal Large Language Models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermined by {hallucination snowballing}: a phenomenon where initial errors amplify across conversational turns, leading to a collapse in coherence. This f…

Cited by 0SourceScholar
2026

OT-ALD: Aligning Latent Distributions with Optimal Transport for Accelerated Image-to-Image Translation

AAAI 2026technical

The Dual Diffusion Implicit Bridge (DDIB) is an emerging image-to-image (I2I) translation method that preserves cycle consistency while achieving strong flexibility. It links two independently trained diffusion models (DMs) in the source and target domains by first adding noise to a source image to

Cited by 0SourcePDFScholar
2025

Can LVLMs Obtain a Driver’s License? A Benchmark Towards Reliable AGI for Autonomous Driving

AAAI 2025technical

Large Vision-Language Models (LVLMs) have recently garnered significant attention, with many efforts aimed at harnessing their general knowledge to enhance the interpretability and robustness of autonomous driving models. However, LVLMs typically rely on large, general-purpose datasets and lack the…

Cited by 4SourcePDFScholar
2025

Mitigating Message Imbalance in Fraud Detection with Dual-View Graph Representation Learning

IJCAI 2025

Graph representation learning has become a mainstream method for fraud detection due to its strong expressive power, which focuses on enhancing node representations through improved neighborhood knowledge capture. However, the focus on local interactions leads to imbalanced transmission of global to

Cited by 0SourcePDFScholar
2025

ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving

ICCV 2025poster

End-to-end autonomous driving has emerged as a promising approach to unify perception, prediction, and planning within a single framework, reducing information loss and improving adaptability. However, existing methods often rely on fixed and sparse trajectory supervision, limiting their ability to…

Cited by 0SourcePDFScholar
2025

Renderworld: World Model with Self-Supervised 3D Label

ICRA 2025

End-to-end autonomous driving with vision-only is not only more cost-effective compared to LiDAR-vision fusion but also more reliable than traditional methods. To achieve a economical and robust purely visual autonomous driving system, we propose RenderWorld, a vision-only end-to-end autonomous driv

Cited by 47SourceScholar
2024

OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree Queries

NeurIPS 2024poster

Occupancy prediction has increasingly garnered attention in recent years for its fine-grained understanding of 3D scenes. Traditional approaches typically rely on dense, regular grid representations, which often leads to excessive computational demands and a loss of spatial details for small objects…

2023

See More and Know More: Zero-shot Point Cloud Segmentation via Multi-modal Visual Data

ICCV 2023poster

Zero-shot point cloud segmentation aims to make deep models capable of recognizing novel objects in point cloud that are unseen in the training phase. Recent trends favor the pipeline which transfers knowledge from seen classes with labels to unseen classes without labels. They typically align visua…

Cited by 33PDFScholar
2023

Semi-supervised Deep Large-Baseline Homography Estimation with Progressive Equivalence Constraint

AAAI 2023technical

Homography estimation is erroneous in the case of large-baseline due to the low image overlay and limited receptive field. To address it, we propose a progressive estimation strategy by converting large-baseline homography into multiple intermediate ones, cumulatively multiplying these intermediate…

2022

Style Mixing and Patchwise Prototypical Matching for One-Shot Unsupervised Domain Adaptive Semantic Segmentation

AAAI 2022technical

In this paper, we tackle the problem of one-shot unsupervised domain adaptation (OSUDA) for semantic segmentation where the segmentors only see one unlabeled target image during training. In this case, traditional unsupervised domain adaptation models usually fail since they cannot adapt to the targ…

2022

Unsupervised Homography Estimation With Coplanarity-Aware GAN

CVPR 2022poster

Estimating homography from an image pair is a fundamental problem in image alignment. Unsupervised learning methods have received increasing attention in this field due to their promising performance and label-free training. However, existing methods do not explicitly consider the problem of plane i…

Cited by 54PDFcodeScholar
2020

Structured Landmark Detection via Topology-Adapting Deep Graph Learning

ECCV 2020poster

Image landmark detection aims to automatically identify the locations of predefined fiducial points. Despite recent success in this field, higher-ordered structural modeling to capture implicit or explicit relationships among anatomical landmarks has not been adequately exploited. In this work, we p…

Cited by 121SourcePDFScholar