2025
LLM Layers Immediately Correct Each Other
NeurIPS 2025poster
Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linear, semantically-meaningful features. Our work demonstrates that an underlying assumption of these methods—that residual stream contributions build addi…