2025
Adaptive Transformer Programs: Bridging the Gap Between Performance and Interpretability in Transformers
ICLR 2025poster
Balancing high performance with interpretability in increasingly powerful Transformer-based models remains a challenge. While mechanistic interpretability aims to specify neural network computations in explicit, pseudocode-like formats, existing methods often involve laborious manual analysis or str…