2026
Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
ICLR 2026poster
Fine-tuned Large Language Models (LLMs) encode rich task-specific features, but the form of these representations—especially within MLP layers—remains unclear. Empirical inspection of LoRA updates shows that new features concentrate in mid-layer MLPs, yet the scale of these layers obscures meaningfu…