← Search

Nirmalendu Prakash

3 accepted papers

2026

Beyond I’m Sorry, I Can’t: Dissecting Large-Language-Model Refusal

AAAI 2026technical

Refusal on harmful prompts is a key safety behaviour in instruction‑tuned large language models (LLMs), yet the internal causes of this behaviour remain poorly understood. We study two public instruction tuned models—Gemma‑2-2B‑IT and LLaMA‑3.1-8B‑IT using sparse autoencoders (SAEs) trained on resid

Cited by 0SourcePDFScholar
2025

Activation Space Interventions Can Be Transferred Between Large Language Models

ICML 2025poster

The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representation universality remain largely unexplored. We bridge this gap by demonstrating that safety interventions can be transfer…

2025

Understanding Refusal in Language Models with Sparse Autoencoders

EMNLP 2025

Refusal is a key safety behavior in aligned language models, yet the internal mechanisms driving refusals remain opaque. In this work, we conduct a mechanistic study of refusal in instruction-tuned LLMs using sparse autoencoders to identify latent features that causally mediate refusal behaviors. We