← Search

Pedram Zaree

1 accepted papers

2025

Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment

EMNLP 2025

Recent research has shown that carefully crafted jailbreak inputs can induce large language models to produce harmful outputs, despite safety measures such as alignment. It is important to anticipate the range of potential Jailbreak attacks to guide effective defenses and accurate assessment of mode

Cited by 0SourcePDFScholar