2025
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
NeurIPS 2025poster
Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large language models (LLMs) for natural language processing (NLP), current automated neuron-level feature description methods face…