2026
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
ICML 2026poster
In mechanistic interpretability, recent work scrutinizes transformer “circuits”—sparse, mono or multi layer sub computations, that may reflect human understandable functions. Yet, these network circuits are rarely acid-tested for their stability across different instances of the same deep learning a…