← Search

Neil Zhenqiang Gong

32 accepted papers

2026

A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent Systems

ICLR 2026poster

Multi-agent systems (MAS) built on large language models (LLMs) increasingly rely on agent-to-agent (A2A) protocols to enable capability discovery, task orchestration, and artifact exchange across heterogeneous stacks. While these protocols promise interoperability, they also introduce new vulnerabi…

Cited by 0SourcecodeScholar
2026

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

ICLR 2026poster

Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field. However, existing benchmarks fail to differentiate question difficulty, limiting their ability…

Cited by 0SourcecodeScholar
2026

TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models

ICLR 2026poster

Generative foundation models (GenFMs), such as large language models and text-to-image systems, have demonstrated remarkable capabilities in various downstream applications. As they are increasingly deployed in high-stakes applications, assessing their trustworthiness has become both a critical nece…

Cited by 0SourceScholar
2026

Watermark-based Attribution of AI-Generated Content

ICLR 2026poster

Several companies have deployed watermark-based detection to identify AI-generated content. However, attribution--the ability to trace back to the user of a generative AI (GenAI) service who created the given AI-generated content--remains largely unexplored despite its growing importance. In this wo…

Cited by 0SourceScholar
2025

BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models

CVPR 2025poster

Multi-modal large language models (MLLMs) extend large language models (LLMs) to process multi-modal information, enabling them to generate responses to image-text inputs. MLLMs have been incorporated into diverse multi-modal applications, such as autonomous driving and medical diagnosis, via plug-a…

Cited by 3SourcePDFScholar
2025

Model Poisoning Attacks to Federated Learning via Multi-Round Consistency

CVPR 2025poster

Model poisoning attacks are critical security threats to Federated Learning (FL). Existing model poisoning attacks suffer from two key limitations: 1) they achieve suboptimal effectiveness when defenses are deployed, and/or 2) they require knowledge of the model updates or local training data on gen…

2025

StringLLM: Understanding the String Processing Capability of Large Language Models

ICLR 2025poster

String processing, which mainly involves the analysis and manipulation of strings, is a fundamental component of modern computing. Despite the significant advancements of large language models (LLMs) in various natural language processing (NLP) tasks, their capability in string processing remains un…

2025

Tracing Back the Malicious Clients in Poisoning Attacks to Federated Learning

NeurIPS 2025poster

Poisoning attacks compromise the training phase of federated learning (FL) such that the learned global model misclassifies attacker-chosen inputs called target inputs. Existing defenses mainly focus on protecting the training phase of FL such that the learnt global model is poison free. However, t…

Cited by 0SourceScholar
2025

WebInject: Prompt Injection Attack to Web Agents

EMNLP 2025

Multi-modal large language model (MLLM)-based web agents interact with webpage environments by generating actions based on screenshots of the webpages. In this work, we propose WebInject, a prompt injection attack that manipulates the webpage environment to induce a web agent to perform an attacker-

Cited by 0SourcePDFScholar
2024

AudioMarkBench: Benchmarking Robustness of Audio Watermarking

NeurIPS 2024poster

The increasing realism of synthetic speech, driven by advancements in text-to-speech models, raises ethical concerns regarding impersonation and disinformation. Audio watermarking offers a promising solution via embedding human-imperceptible watermarks into AI-generated audios. However, the robustne…

2024

Certifiably Robust Image Watermark

ECCV 2024poster

"Generative AI raises many societal concerns such as boosting disinformation and propaganda campaigns. Watermarking AI-generated content is a key technology to address these concerns and has been widely deployed in industry. However, watermarking is vulnerable to removal attacks and forgery attacks.…

2024

Data Poisoning based Backdoor Attacks to Contrastive Learning

CVPR 2024poster

Contrastive learning (CL) pre-trains general-purpose encoders using an unlabeled pre-training dataset which consists of images or image-text pairs. CL is vulnerable to data poisoning based backdoor attacks (DPBAs) in which an attacker injects poisoned inputs into the pre-training dataset so the enco…

2024

DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

ICLR 2024spotlight

Large language models (LLMs) have achieved remarkable performance in various evaluation benchmarks. However, concerns are raised about potential data contamination in their considerable volume of training corpus. Moreover, the static nature and fixed complexity of current benchmarks may inadequately…

2024

FedREDefense: Defending against Model Poisoning Attacks for Federated Learning using Model Update Reconstruction Error

ICML 2024poster

Federated Learning (FL) faces threats from model poisoning attacks. Existing defenses, typically relying on cross-client/global information to mitigate these attacks, fall short when faced with non-IID data distributions and/or a large number of malicious clients. To address these challenges, we pre…

Cited by 7SourcePDFScholar
2024

MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

ICLR 2024poster

Large language models (LLMs) have garnered significant attention due to their impressive natural language processing (NLP) capabilities. Recently, many studies have focused on the tool utilization ability of LLMs. They primarily investigated how LLMs effectively collaborate with given specific tools…

2024

Position: TrustLLM: Trustworthiness in Large Language Models

ICML 2024poster

Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLM…

Cited by 95SourcePDFScholar
2024

ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods

EMNLP 2024main

The rapid scaling of large language models (LLMs) has raised concerns about the transparency and fair use of the data used in their pretraining. Detecting such content is challenging due to the scale of the data and limited exposure of each instance during training. We propose ReCaLL (Relative Condi…

Cited by 11SourcePDFScholar
2024

Unlocking the Potential of Federated Learning: The Symphony of Dataset Distillation via Deep Generative Latents

ECCV 2024poster

"Data heterogeneity presents significant challenges for federated learning (FL). Recently, dataset distillation techniques have been introduced, and performed at the client level, to attempt to mitigate some of these challenges. In this paper, we propose a highly efficient FL dataset distillation fr…

2023

PointCert: Point Cloud Classification With Deterministic Certified Robustness Guarantees

CVPR 2023poster

Point cloud classification is an essential component in many security-critical applications such as autonomous driving and augmented reality. However, point cloud classifiers are vulnerable to adversarially perturbed point clouds. Existing certified defenses against adversarial point clouds suffer f…

Cited by 11SourcePDFScholar
2022

Addressing Heterogeneity in Federated Learning via Distributional Transformation

ECCV 2022poster

"Federated learning (FL) allows multiple clients to collaboratively train a deep learning model. One major challenge of FL is when data distribution is heterogeneous, i.e., differs from one client to another. Existing personalized FL algorithms are only applicable to narrow cases, e.g., one or two d…

2022

Almost Tight L0-norm Certified Robustness of Top-k Predictions against Adversarial Perturbations

ICLR 2022poster

Top-$k$ predictions are used in many real-world applications such as machine learning as a service, recommender systems, and web searches. $\ell_0$-norm adversarial perturbation characterizes an attack that arbitrarily modifies some features of an input such that a classifier makes an incorrect pred…

Cited by 26SourcePDFScholar
2022

Certified Robustness of Nearest Neighbors against Data Poisoning and Backdoor Attacks

AAAI 2022technical

Data poisoning attacks and backdoor attacks aim to corrupt a machine learning classifier via modifying, adding, and/or removing some carefully selected training examples, such that the corrupted classifier makes incorrect predictions as the attacker desires. The key idea of state-of-the-art certifie…

2022

MultiGuard: Provably Robust Multi-label Classification against Adversarial Examples

NeurIPS 2022accept

Multi-label classification, which predicts a set of labels for an input, has many applications. However, multiple recent studies showed that multi-label classification is vulnerable to adversarial examples. In particular, an attacker can manipulate the labels predicted by a multi-label classifier f…

2022

Semi-Leak: Membership Inference Attacks against Semi-Supervised Learning

ECCV 2022poster

"Semi-supervised learning (SSL) leverages both labeled and unlabeled data to train machine learning (ML) models. State-of-the-art SSL methods can achieve comparable performance to supervised learning by leveraging much fewer labeled data. However, most existing works focus on improving the performan…

2021

Intrinsic Certified Robustness of Bagging against Data Poisoning Attacks

AAAI 2021technical

In a data poisoning attack, an attacker modifies, deletes, and/or inserts some training examples to corrupt the learnt machine learning model. Bootstrap Aggregating (bagging) is a well known ensemble learning method, which trains multiple base models on random subsamples of a training dataset using…

2021

Semi-Supervised Node Classification on Graphs: Markov Random Fields vs. Graph Neural Networks

AAAI 2021technical

Semi-supervised node classification on graph-structured data has many applications such as fraud detection, fake account and review detection, user’s private attribute inference in social networks, and community detection. Various methods such as pairwise Markov Random Fields (pMRF) and graph neura…

Cited by 26SourcePDFScholar
2020

Certified Robustness for Top-k Predictions against Adversarial Perturbations via Randomized Smoothing

ICLR 2020poster

It is well-known that classifiers are vulnerable to adversarial perturbations. To defend against adversarial perturbations, various certified robustness results have been derived. However, existing certified robustnesses are limited to top-1 predictions. In many real-world applications, top-$k$ pre…

Cited by 108SourcecodeScholar