← Search

Siyuan Fu

1 accepted papers

2026

SafeSeek: Universal Attribution of Safety Circuits in Language Models

ICML 2026poster

Mechanistic interpretability reveals that safety-critical behaviors (e.g., alignment, jailbreak, backdoor) in Large Language Models (LLMs) are grounded in specialized functional components. However, existing safety attribution methods struggle with generalization and reliability due to their relianc…

Cited by 0SourceScholar