← Search

Mengxuan Wang

2 accepted papers

2026

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

ICML 2026poster

Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerable to multimodal jailbreak attacks. Existing defenses predominantly rely on safety fine-tuning or \textit{aggressive} token manipulations, incurring subs…

Cited by 0SourceScholar
2025

Less Yet Robust: Crucial Region Selection for Scene Recognition

ICASSP 2025accepted

Scene recognition, particularly for aerial and underwater images, often suffers from various types of degradation, such as blurring or overexposure. Previous works that focus on convolutional neural networks have been shown to be able to extract panoramic semantic features and perform well on scene…

Cited by 0SourceScholar