← Search

Wenqian Ye

12 accepted papers

2026

SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias

AAAI 2026technical

Large vision-language models such as CLIP have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models often develop multimodal spurious biases, the undesirable tendency to rely on spurious features. For example, CLIP may infer

Cited by 0SourcePDFScholar
2025

On-Board Vision-Language Models (VLMs) for Personalized Motion Control of Autonomous Vehicles

IROS 2025

Personalized driving refers to an autonomous vehicle’s ability to adapt its driving behavior or control strategies to match individual users’ preferences and driving styles while maintaining safety and comfort standards. However, existing works either fail to capture every individual’s preference pr

Cited by 1SourceScholar
2024

LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs

CVPR 2024poster

Autonomous driving (AD) has made significant strides in recent years. However existing frameworks struggle to interpret and execute spontaneous user instructions such as "overtake the car ahead." Large Language Models (LLMs) have demonstrated impressive reasoning capabilities showing potential to br…

2024

Learning Autonomous Driving Tasks via Human Feedbacks with Large Language Models

EMNLP 2024finding

Traditional autonomous driving systems have mainly focused on making driving decisions without human interaction, overlooking human-like decision-making and human preference required in complex traffic scenarios. To bridge this gap, we introduce a novel framework leveraging Large Language Models (LL…

Cited by 2SourcePDFScholar
2024

Learning Robust Classifiers with Self-Guided Spurious Correlation Mitigation

IJCAI 2024poster

Deep neural classifiers tend to rely on spurious correlations between spurious attributes of inputs and targets to make predictions, which could jeopardize their generalization capability. Training classifiers robust to spurious correlations typically relies on annotations of spurious correlations i…

2024

MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding

CVPR 2024poster

Vision-language generative AI has demonstrated remarkable promise for empowering cross-modal scene understanding of autonomous driving and high-definition (HD) map systems. However current benchmark datasets lack multi-modal point cloud image and language data pairs. Recent approaches utilize visual…

2023

Mitigating Transformer Overconfidence via Lipschitz Regularization

UAI 2023poster

Though Transformers have achieved promising results in many computer vision tasks, they tend to be over-confident in predictions, as the standard Dot Product Self-Attention (DPSA) can barely preserve distance for the unbounded input domain. In this work, we fill this gap by proposing a novel Lipschi…

2023

Vitasd: Robust Vision Transformer Baselines for Autism Spectrum Disorder Facial Diagnosis

ICASSP 2023accepted

Autism spectrum disorder (ASD) is a lifelong neurodevelopmental disorder with very high prevalence around the world. Research progress in the field of ASD facial analysis in pediatric patients has been hindered due to a lack of well-established baselines. In this paper, we propose the use of the Vis…

Cited by 0SourceScholar