2025
COSMIC: Generalized Refusal Direction Identification in LLM Activations
ACL 2025finding
Large Language Models encode behaviors like refusal within their activation space, but identifying these behaviors remains challenging. Existing methods depend on predefined refusal templates detectable in output tokens or manual review. We introduce **COSMIC** (Cosine Similarity Metrics for Inversi…