← Search

Zhaoquan Gu

10 accepted papers

2026

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models

ICML 2026poster

Multimodal large language models (MLLMs) inevitably memorize sensitive cross-modal information during pretraining, making post-deployment unlearning crucial for safety. Existing methods often evaluate unlearning based on output deviations, neglecting generation quality, which can lead to hallucinati…

Cited by 0SourceScholar
2026

Debiased Dual-Invariant Defense for Adversarially Robust Person Re-Identification

AAAI 2026technical

Person re-identification (ReID) is a fundamental task in many real-world applications such as pedestrian trajectory tracking. However, advanced deep learning-based ReID models are highly susceptible to adversarial attacks, where imperceptible perturbations to pedestrian images can cause entirely inc

Cited by 0SourcePDFScholar
2026

Nasty Adversarial Training: A Probability Sparsity Perspective for Robustness Enhancement

ICLR 2026poster

The vulnerability of deep neural networks to adversarial examples poses significant challenges to their reliable deployment. Among existing empirical defenses, adversarial training and robust distillation have proven the most effective. In this paper, we identify a property originally associated wit…

Cited by 0SourceScholar
2026

Toward Subspace-Perturbed Trajectory-Aware Backdoor Attacks in Deep Reinforcement Learning

ICML 2026poster

Deep Reinforcement Learning agents are in- creasingly used in safety-critical domains but remain vulnerable to stealthy backdoor attacks. Existing outer-loop attacks face a trade-off be- tween perceptual stealth, poisoning efficiency, and value-function consistency, often making the at- tack ineffec…

Cited by 0SourceScholar
2025

Fast Think-on-Graph: Wider, Deeper and Faster Reasoning of Large Language Model on Knowledge Graph

AAAI 2025technical

Graph Retrieval Augmented Generation (GRAG) is a novel paradigm that takes the naive RAG system a step further by integrating graph information, such as knowledge graph (KGs), into large-scale language models (LLMs) to mitigate hallucination. However, existing GRAG still encounter limitations: 1) si…

2023

Deep Manifold Attack on Point Clouds via Parameter Plane Stretching

AAAI 2023technical

Adversarial attack on point clouds plays a vital role in evaluating and improving the adversarial robustness of 3D deep learning models. Current attack methods are mainly applied by point perturbation in a non-manifold manner. In this paper, we formulate a novel manifold attack, which deforms the un…

Cited by 17SourcePDFScholar
2022

Filter Pruning via Feature Discrimination in Deep Neural Networks

ECCV 2022poster

"Filter pruning is one of the most effective methods to compress deep convolutional networks (CNNs). In this paper, as a key component in filter pruning, We first propose a feature discrimination based filter importance criterion, namely Receptive Field Criterion (RFC). It turns the maximum activati…

Cited by 28SourcePDFScholar
2022

Improving Robustness of Language Models from a Geometry-aware Perspective

ACL 2022findings

Recent studies have found that removing the norm-bounded projection and increasing search steps in adversarial training can significantly improve robustness. However, we observe that a too large number of search steps can hurt accuracy. We aim to obtain strong robustness efficiently using fewer step…

2022

Robust Network Architecture Search via Feature Distortion Restraining

ECCV 2022poster

"The vulnerability of DNNs severely limits the application of it in the security-sensitive domains. Most of the existing methods improve the robustness of models from weight optimization, such as adversarial training and regularization. However, the architecture is also a key factor to robustness, w…

Cited by 7SourcePDFScholar
2021

CODEs: Chamfer Out-of-Distribution Examples Against Overconfidence Issue

ICCV 2021poster

Overconfident predictions on out-of-distribution (OOD) samples is a thorny issue for deep neural networks. The key to resolve the OOD overconfidence issue inherently is to build a subset of OOD samples and then suppress predictions on them. This paper proposes the Chamfer OOD examples (CODEs), whose…

Cited by 39PDFScholar