2025
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
ICML 2025poster
Interpretability research now offers a variety of techniques for identifying abstract internal mechanisms in neural networks. Can such techniques be used to predict how models will behave on out-of-distribution examples? In this work, we provide a positive answer to this question. Through a diverse…