← Search

Narmeen Fatimah Oozeer

3 accepted papers

2025

Activation Space Interventions Can Be Transferred Between Large Language Models

ICML 2025poster

The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representation universality remain largely unexplored. We bridge this gap by demonstrating that safety interventions can be transfer…

2025

Beyond Linear Steering: Unified Multi-Attribute Control for Language Models

EMNLP 2025

Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of linear steering methods, which assume additive behavior in activation space and require per-attribute tuning. We introdu

Cited by 0SourcePDFScholar
2025

Position: Require Frontier AI Labs To Release Small "Analog" Models

NeurIPS 2025poster

Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovation tradeoff. This paper argues for an alternative regulatory approach that ensures AI safety while actively \textit{prom…

Cited by 0SourceScholar