← Search

Yannan Liu

3 accepted papers

2025

DF-MIA: A Distribution-Free Membership Inference Attack on Fine-Tuned Large Language Models

AAAI 2025technical

Membership Inference Attack (MIA) aims to determine if a specific sample is present in the training dataset of a target machine learning model. Previous MIAs against fine-tuned Large Language Models (LLMs) either fail to address the unique challenges in the fine-tuned setting or rely on strong assu…

2025

One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs

ICLR 2025poster

Safety alignment in large language models (LLMs) is increasingly compromised by jailbreak attacks, which can manipulate these models to generate harmful or unintended content. Investigating these attacks is crucial for uncovering model vulnerabilities. However, many existing jailbreak strategies fai…

2021

TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning Tasks

NeurIPS 2021poster

Deep learning (DL) systems are notoriously difficult to test and debug due to the lack of correctness proof and the huge test input space to cover. Given the ubiquitous unlabeled test data and high labeling cost, in this paper, we propose a novel test prioritization technique, namely TestRank, which…

Cited by 31SourcePDFScholar