← Search

Shanmin Pang

8 accepted papers

2026

PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems

AAAI 2026technical

Multimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scen

Cited by 0SourcePDFScholar
2026

Redundant Queries in DETR-Based 3D Detection Methods: Unnecessary and Prunable

AAAI 2026technical

Query-based models are extensively used in 3D object detection tasks, with a wide range of pre-trained checkpoints readily available online. However, despite their popularity, these models often require an excessive number of object queries, far surpassing the actual number of objects to detect. The

Cited by 0SourcePDFScholar
2025

Accelerate 3D Object Detection Models via Zero-Shot Attention Key Pruning

ICCV 2025poster

Query-based methods with dense features have demonstrated remarkable success in 3D object detection tasks. However, the computational demands of these models, particularly with large image sizes and multiple transformer layers, pose significant challenges for efficient running on edge devices. Exist…

2025

Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis

ICCV 2025poster

Egocentricly comprehending the causes and effects of car accidents is crucial for the safety of self-driving cars, and synthesizing causal-entity reflected accident videos can facilitate the capability test to respond to unaffordable accidents in reality. However, incorporating causal relations as s…

Cited by 0SourcePDFScholar
2024

ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep Generation

NeurIPS 2024poster

The commercial text-to-image deep generation models (e.g. DALL·E) can produce high-quality images based on input language descriptions. These models incorporate a black-box safety filter to prevent the generation of unsafe or unethical content, such as violent, criminal, or hateful imagery. Recent j…

Cited by 2SourcePDFScholar
2024

TICMapNet: A Tightly Coupled Temporal Fusion Pipeline for Vectorized HD Map Learning

RA-L 2024

High-Definition (HD) map construction is essential for autonomous driving to accurately understand the surrounding environment. Most existing methods rely on single-frame inputs to predict local map, which often fail to effectively capture the temporal correlations between frames. This limitation re

Cited by 7SourceScholar
2022

Sparse Semantic Map-Based Monocular Localization in Traffic Scenes Using Learned 2D-3D Point-Line Correspondences

RA-L 2022

Vision-based localization in a prior map is of crucial importance for autonomous vehicles. Given a query image, the goal is to estimate the camera pose corresponding to the prior map, and the key is the registration problem of camera images within the map. While autonomous vehicles drive on the road

Cited by 9SourceScholar
2021

MagDR: Mask-Guided Detection and Reconstruction for Defending Deepfakes

CVPR 2021poster

Deepfakes raised serious concerns on the authenticity of visual contents. Prior works revealed the possibility to disrupt deepfakes by adding adversarial perturbations to the source data, but we argue that the threat has not been eliminated yet. This paper presents MagDR, a mask-guided detection and…

Cited by 44PDFScholar