← Search

Jinbao Li

8 accepted papers

2026

From ``Sure" to ``Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons

ICLR 2026poster

Large Vision-Language Models (LVLMs) are vulnerable to jailbreak attacks that can generate harmful content. Existing detection methods are either limited to detecting specific attack types or are too time-consuming, making them impractical for real-world deployment. To address these challenges, we p…

Cited by 0SourcecodeScholar
2025

CAMH: Advancing Model Hijacking Attack in Machine Learning

AAAI 2025technical

In the burgeoning domain of machine learning, the reliance on third-party services for model training and the adoption of pre-trained models have surged. However, this reliance introduces vulnerabilities to model hijacking attacks, where adversaries manipulate models to perform unintended tasks, lea…

Cited by 0SourcePDFScholar
2025

Enhancing Adversarial Transferability with Adversarial Weight Tuning

AAAI 2025technical

Deep neural networks (DNNs) are vulnerable to adversarial examples (AEs) that mislead the model while appearing benign to human observers. A critical concern is the transferability of AEs, which enables black-box attacks without direct access to the target model. However, many previous attacks have…

Cited by 0SourcePDFScholar
2025

TWIST: Text-encoder Weight-editing for Inserting Secret Trojans in Text-to-Image Models

ACL 2025long

Text-to-image (T2I) models excel at generating high-quality images from text via powerful text encoders but training these encoders demands substantial computational resources. Consequently, many users seek pre-trained text encoders from model plugin-sharing platforms like Civitai and Hugging Face,…

Cited by 0SourcePDFScholar
2024

Integer Is Enough: When Vertical Federated Learning Meets Rounding

AAAI 2024technical

Vertical Federated Learning (VFL) is a solution increasingly used by companies with the same user group but differing features, enabling them to collaboratively train a machine learning model. VFL ensures that clients exchange intermediate results extracted by their local models, without sharing ra…

Cited by 2SourcePDFScholar
2024

Protecting Object Detection Models from Model Extraction Attack via Feature Space Coverage

IJCAI 2024poster

The model extraction attack is an attack pattern aimed at stealing well-trained machine learning models' functionality or privacy information. With the gradual popularization of AI-related technologies in daily life, various well-trained models are being deployed. As a result, these models are consi…

2021

Attention-Embedded Decomposed Network with Unpaired CT Images Prior for Metal Artifact Reduction

ICASSP 2021accepted

Recently, unsupervised learning is proposed to avoid the performance degrading caused by synthesized paired computed tomography (CT) images. However, existing unsupervised methods for metal artifact reduction (MAR) only use features in image space, which is not enough to restore regions heavily corr…

Cited by 0SourceScholar