2024
Compact Proofs of Model Performance via Mechanistic Interpretability
NeurIPS 2024poster
We propose using mechanistic interpretability -- techniques for reverse engineering model weights into human-interpretable algorithms -- to derive and compactly prove formal guarantees on model performance. We prototype this approach by formally proving accuracy lower bounds for a small transformer…