← Search

Neil Gong

3 accepted papers

2026

CREDIT: Certified Ownership Verification of Deep Neural Networks Against Model Extraction Attacks

ICML 2026poster

Machine Learning as a Service (MLaaS) has become a widely adopted method for delivering deep neural network (DNN) models, allowing users to conveniently access models via APIs. However, such services have been shown to be highly vulnerable to Model Extraction Attacks (MEAs). While numerous defense s…

Cited by 0SourceScholar
2024

GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis

ACL 2024long

Large Language Models (LLMs) face threats from jailbreak prompts. Existing methods for detecting jailbreak prompts are primarily online moderation APIs or finetuned LLMs. These strategies, however, often require extensive and resource-intensive data collection and training processes. In this study,…

2024

Visual Hallucinations of Multi-modal Large Language Models

ACL 2024findings

Visual hallucination (VH) means that a multi-modal LLM (MLLM) imagines incorrect details about an image in visual question answering. Existing studies find VH instances only in existing image datasets, which results in biased understanding of MLLMs’ performance under VH due to limited diversity of s…