← Search

Peijie Sun

4 accepted papers

2026

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

ICML 2026poster

The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constraints on parameters, gradients, or internal representations, we observe that they can be effectively circumvented under persistent HFT. Our analysis traces this …

Cited by 0SourceScholar
2024

A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models

EMNLP 2024main

Large language models (LLMs) are essential tools that users employ across various scenarios, so evaluating their performance and guiding users in selecting the suitable service is important. Although many benchmarks exist, they mainly focus on specific predefined model abilities, such as world knowl…