← Search

Amir Abdullah

6 accepted papers

2026

Beyond I’m Sorry, I Can’t: Dissecting Large-Language-Model Refusal

AAAI 2026technical

Refusal on harmful prompts is a key safety behaviour in instruction‑tuned large language models (LLMs), yet the internal causes of this behaviour remain poorly understood. We study two public instruction tuned models—Gemma‑2-2B‑IT and LLaMA‑3.1-8B‑IT using sparse autoencoders (SAEs) trained on resid

Cited by 0SourcePDFScholar
2025

Activation Space Interventions Can Be Transferred Between Large Language Models

ICML 2025poster

The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representation universality remain largely unexplored. We bridge this gap by demonstrating that safety interventions can be transfer…

2025

Beyond Linear Steering: Unified Multi-Attribute Control for Language Models

EMNLP 2025

Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of linear steering methods, which assume additive behavior in activation space and require per-attribute tuning. We introdu

Cited by 0SourcePDFScholar
2025

TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research

EMNLP 2025

Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models. To bridge this gap, we propose text-to-SQL generation as an ideal task to study, as it combines the formal structure of toy tasks with real-world complexity. We

Cited by 0SourcePDFScholar
2024

Interpreting Learned Feedback Patterns in Large Language Models

NeurIPS 2024poster

Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferences in human feedback data. We coin the term **Learned Feedback Pattern** (LFP) for patterns in an LLM's activations lea…

2023

PCMID: Multi-Intent Detection through Supervised Prototypical Contrastive Learning

EMNLP 2023long findings

Intent detection is a major task in Natural Language Understanding (NLU) and is the component of dialogue systems for interpreting users’ intentions based on their utterances. Many works have explored detecting intents by assuming that each utterance represents only a single intent. Such systems hav…

Cited by 0SourceScholar