← Search

Maithili Joshi

1 accepted papers

2025

SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection

EMNLP 2025

Large Language Models (LLMs) with safe-alignment training are powerful instruments with robust language comprehension capability. Typically LLMs undergo careful alignment training involving human feedback to ensure the acceptance of safe inputs while rejection of harmful or unsafe ones. However, the