← Search

Ruben Härle

2 accepted papers

2026

ActivationReasoning: Logical Reasoning in Latent Activation Spaces

ICLR 2026poster

Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control. Sparse autoencoders (SAEs) make hidden activations more interpretable by exposing latent features that often align with human concepts. Yet, these features are fragile…

Cited by 0SourcecodeScholar
2025

Measuring and Guiding Monosemanticity

NeurIPS 2025spotlight

There is growing interest in leveraging mechanistic interpretability and controllability to better understand and influence the internal dynamics of large language models (LLMs). However, current methods face fundamental challenges in reliably localizing and manipulating feature representations. Spa…

Cited by 0SourceScholar