Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders
Sparse Autoencoders (SAEs) are widely used to steer large language models (LLMs), based on the assumption that their interpretable features naturally enable effective model behavior steering. Yet a fundamental question remains: does higher interpretability imply better steering utility? To answer th…