2026
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
ICLR 2026poster
Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control. Sparse autoencoders (SAEs) make hidden activations more interpretable by exposing latent features that often align with human concepts. Yet, these features are fragile…