← Search

Zhengming Zhang

7 accepted papers

2025

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

ICCV 2025poster

Despite the remarkable performance of multimodal large language models (MLLMs) across diverse tasks, the substantial training and inference costs impede their advancement. In this paper, we propose p-MoD, an efficient MLLM architecture that significantly reduces training and inference costs while ma…

2024

SceneDiff: Generative Scene-Level Image Retrieval with Text and Sketch Using Diffusion Models

IJCAI 2024poster

Jointly using text and sketch for scene-level image retrieval utilizes the complementary between text and sketch to describe the fine-grained scene content and retrieve the target image, which plays a pivotal role in accurate image retrieval. Existing methods directly fuse the features of sketch and…

Cited by 0SourcePDFScholar
2024

Teach LLMs to Phish: Stealing Private Information from Language Models

ICLR 2024poster

When large language models are trained on private data, it can be a \textit{significant} privacy risk for them to memorize and regurgitate sensitive information. In this work, we propose a new \emph{practical} data extraction attack that we call ``neural phishing''. This attack enables an adversary…

Cited by 27SourcePDFScholar
2023

AIRA-DA: Adversarial Image Reconstruction Alignments for Unsupervised Domain Adaptive Object Detection

RA-L 2023

Unsupervised domain adaptive object detection is a challenging perception task where object detectors are adapted from a label-rich source domain to an unlabeled target domain, playing a vital role in autonomous driving and robot navigation. Since the camera settings, weather, and light conditions v

Cited by 7SourceScholar
2023

TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty

AAAI 2023technical

With rapid development in hardware (sensors and processors) and AI algorithms, automated driving techniques have entered the public’s daily life and achieved great success in supporting human driving performance. However, due to the high contextual variations and temporal dynamics in pedestrian beha…

2022

Neurotoxin: Durable Backdoors in Federated Learning

ICML 2022spotlight

Federated learning (FL) systems have an inherent vulnerability to adversarial backdoor attacks during training due to their decentralized nature. The goal of the attacker is to implant backdoors in the learned model with poisoned updates such that at test time, the model’s outputs can be fixed to a…

2021

ECS-Net: Improving Weakly Supervised Semantic Segmentation by Using Connections Between Class Activation Maps

ICCV 2021poster

Image-level weakly supervised semantic segmentation is a challenging task. As classification networks tend to capture notable object features and are insensitive to overactivation, class activation map (CAM) is too sparse and rough to guide segmentation network training. Inspired by the fact that er…

Cited by 137PDFScholar